Why this quant
- 🗜️ 3.2× smaller than BF16, and smaller than every official quant. 19.3 GiB against 62.3 GiB BF16, 33.0 GiB FP8, 22.4 GiB INT4 and 20.1 GiB NVFP4.
- 🎯 Accuracy is tied with BF16. 81.7 against BF16's 81.9, pooled over both suites, inside the ±1-point run-to-run spread. The official FP8 ties too; the official NVFP4 does not.
- ⚡ 2.07× BF16 throughput with DFlash on. 3,268 tok/s at concurrency 32, ahead of the official NVFP4's 3,182 while being 0.8 GiB smaller. Without a drafter that order reverses; both regimes are in the table below.
- 🥇 Smallest of the field, and the most accurate 4-bit-class build of it. 81.7 pooled against the official NVFP4's 79.5, a 2.2-point lead outside the tie band, at 0.8 GiB less on disk.
- 💾 Fits where BF16 cannot. Comfortable on a 96 GiB card and on 40 GiB-class GPUs, where the 62.3 GiB original does not load at all.
- 🔧 No calibration data, no custom runtime. Weights-only round-to-nearest,
compressed-tensors, stock vllm serve.
Serve it
vllm serve primitive-ai/Laguna-XS-2.1-mixed-NVFP4-MXFP8
DFlash speculative decoding works on this checkpoint. The accuracy table below runs without it; the
DFlash throughput numbers are in their own table further down.
Measured
1,370 items across fourteen public benchmarks. A 1,170-item knowledge suite and a 200-item
tool-calling suite, under one fixed protocol for every row: temperature 0.6 / top_p 0.95 /
top_k 20, thinking forced on, a 16,384-token budget, concurrency 32, no speculative decoding, all
rows back to back in one sitting on one RTX PRO 6000 Blackwell.
Table with columns: build, bpw, size, overall, knowledge, call, abstain, runs k/a, finished, out/answer, tok/s @ 32, per-token latency| build | bpw | size | overall | knowledge | call | abstain | runs k/a | finished | out/answer | tok/s @ 32 | per-token latency |
|---|
| BF16 (reference) | 16.000 | 62.3 G | 81.9 | 84.0 | 69.4 | 70.0 | 1/1 |
overall is one number over both suites: the 1,170 knowledge and 200 tool-calling items pooled
as 1,370, weighted 85.4% and 14.6% by item count. Repeat runs of one checkpoint moved it by
about a point, so treat a gap below 1.0 as a tie.
Table with columns: build, agentic, call, abstain| build | agentic | call | abstain |
|---|
| BF16 (reference) | 69.5 | 69.4 | 70.0 |
| official FP8 | 69.5 | 70.6 | 65.0 |
| official NVFP4 | 62.5 | 59.4 | 75.0 |
|
call is accuracy on the 160 rows that require a call; abstain is the 40 whose correct action is
to call nothing. Benchmarks: BFCL v4 (30, including irrelevance), xLAM/APIGen (45), ToolACE
(45), Glaive v2 (40), nvidia When2Call (40, the abstention rows). Tool schemas go in the system prompt
and the model answers with a JSON array of calls, the same way for every model. An item counts only
if every expected call is present with acceptable arguments and no call is invented.
Abstention is the weak axis for every model we have measured (52–82%), so a build can look strong
on overall and still over-call.
‡ poolside's official INT4 no longer loads on our vLLM build. It asserts during engine
initialisation, so it has no accuracy row here. Its speed row below was measured when it did.
Thinking is forced on, and on this model that is the difference between two different
products. Every checkpoint in this family ships a chat template that defaults enable_thinking
to false. Left at that default this build answers in 268 tokens and scores 80.4; made to
reason it writes 1,045 tokens and scores 83.8. Both are real numbers about the same weights, and
the second one is the one you get if you serve it the way the rest of the field is served, so the
protocol pins the flag rather than trusting the template.
Where this build sits. It ties BF16 (81.7 against 81.9) and the official FP8, and leads the
official NVFP4 by 2.2 points, the one gap in this table outside the ±1-point band. That gap is
almost entirely tool calling: 69.4 against 62.5, where the NVFP4 build both calls less accurately and
invents more. On size it is the smallest row here.
On speed, speculation changes the order, so both regimes are here. Without a drafter the official
NVFP4 is the faster build, 16.3 ms per token against 20.7, and it gets there partly by writing
longer answers (1,310 tokens against 1,045), which lifts aggregate throughput at a fixed
concurrency while costing more per answer. With DFlash attached, this build leads instead. If you
serve without a drafter, take the NVFP4 build's latency seriously.
With DFlash speculative decoding on. A separate configuration from the accuracy pass above,
measured in its own sitting with speculation enabled for every row:
Table with columns: build, tok/s @ 32, × BF16| build | tok/s @ 32 | × BF16 |
|---|
| BF16 | 1,579 | 1.00 |
| official FP8 | 917 | 0.58 |
| official INT4 | 1,301 | 0.82 |
| official NVFP4 | 3,182 | 2.01 |
| this repo | 3,268 | 2.07 |
Single-stream decode, measured separately without speculation: 196 tok/s.
Comparable with our other models
Accuracy numbers move for reasons that have nothing to do with the model: a shorter token budget, a
different temperature, or whether the model was allowed to reason at all. So every number in this
table, on this card and on our other cards, comes from one fixed protocol.
The same 1,370 items: a 1,170-item knowledge suite (MMLU-Pro, ARC-Challenge, HellaSwag, WinoGrande,
CommonsenseQA, BoolQ, OpenBookQA, GSM8K, MATH-500) and a 200-item tool-calling suite (BFCL v4,
xLAM/APIGen, ToolACE, Glaive v2, nvidia When2Call). temperature 0.6, top_p 0.95, top_k 20,
thinking forced on, a 16,384-token budget, no reasoning parser, scoring the last ANSWER: in the
reply. Concurrency 32 on one RTX PRO 6000 Blackwell, each model's rows in one sitting. Auto-scored,
no LLM judge. Both halves are means of at least three runs per build.
Table with columns: model, shape, size, overall, knowledge, call, abstain, finished, out, tok/s @ 32| model | shape | size | overall | knowledge | call | abstain | finished | out | tok/s @ 32 |
|---|
| Laguna-XS-2.1 (this repo) | 31 B MoE | 19.3 GiB | 81.7 | 83.8 | 68.4 | 73.5 | 98.9% | 1097 tok | 1523 |
Read overall with finished. overall scores an answer that overran the token budget as wrong,
but it cannot say whether the model needed the room or failed to stop; finished and out separate
those. A gap under 1.0 is a tie. The tok/s column comes from each model's own sitting and drifts
a few percent between sittings, so read it as a bracket.
call and abstain are the tool-calling suite's two halves, reported separately. call is
accuracy on the 160 items that require a tool call; abstain is the 40 whose correct action is to
call nothing. They used to be pooled into one agentic number, and the pooling misled: a model with
ordinary call accuracy and unusual abstention discipline outscored models that are better at
actually making calls. Weight them by your own workload's mix.
Table with columns: benchmark, Laguna-XS-2.1, Nemotron-3.5-Lightning-30B-A3B, Ornith-1.5-35B-A3B, Muse-Glimmer-30B, Qwen3.8-27B, Laguna-S-2.1, Qwen3.8-Flash-Next| benchmark | Laguna-XS-2.1 | Nemotron-3.5-Lightning-30B-A3B | Ornith-1.5-35B-A3B | Muse-Glimmer-30B | Qwen3.8-27B | Laguna-S-2.1 | Qwen3.8-Flash-Next |
|---|
| knowledge | | | | | | | |
| mmlu_pro | 79.0 | 82.0 |
What's quantized to what
Table with columns: tensors, format| tensors | format |
|---|
| the bulk of the MoE expert projections | NVFP4 (group 16) |
| attention projections, shared-expert projections, one expert layer | MXFP8 (group 32) |
embeddings, lm_head, router, attention gates, norms | BF16 |
Weights-only round-to-nearest, no calibration data required or embedded.