Trained purely on skills: real error→diagnosis→fix traces, implementation tasks, and tool-grounded answers — exact facts are delegated to retrieval by design.
Benchmark (40 app-building tasks and 210 knowledge questions the models never saw during training, scored against real, tested reference implementations): 53.2% knowledge retention via tool use; 47.1% on build tasks; tool calibration retained (95–98% / 7.5%). For comparison, Claude Sonnet 5 scored 31.5% on the same build tasks.
- Base:
google/gemma-4-26B-A4B-it (Apache-2.0, subject to Gemma Terms)
- Training: LoRA supervised fine-tuning on the Agent University corpus — ~90 live-tested curricula for AI/agent libraries and dev tools (Supabase, MCP, Next.js, Cloudflare, Slack, and more)
- Format: Fireworks
fused_peft_3d_v1 layout — MoE expert LoRAs are fused 3D tensors. Standard PEFT loads the attention/MLP portion; for a full-fidelity merge use merge_adapter.py from the agent-university harness. A merged, ready-to-run Q8_0 GGUF is in the companion -gguf repo.