Why this quant
- 🖥️ Runs on one 96 GiB GPU. BF16 needs a multi-GPU host, and poolside's own INT4/NVFP4
releases are now published at 92.9 GB, which no longer loads on a 95 GiB card.
- 📦 Joint-smallest of the 4-bit field at 64.0 GiB, level with olka-fi's MXFP4 and 3–8 GiB under
the NVFP4 and AWQ builds.
- ⚡ The fastest at concurrency 32. 670 tok/s against 626–663 for the other three builds,
measured back to back in one sitting.
- 🎯 Accuracy mid-field in a narrow field. 84.3 pooled over both suites against 85.6 for the AWQ
build and 83.5 for the NVFP4 one.
- 🔧 No calibration data. Weights-only round-to-nearest, so the conversion is data-free and
reproducible.
Serve it
vllm serve primitive-ai/Laguna-S-2.1-mixed-NVFP4-MXFP8
Sized for one 96 GiB card at --gpu-memory-utilization 0.90. Give it a generous
--max-model-len: this model reasons at length and a short cap will cut answers off mid-thought.
No sampling workarounds are needed. See the revision note at the bottom if you pulled this repo
before 19 Aug 2026.
Measured
1,370 items across fourteen public benchmarks. A 1,170-item knowledge suite and a 200-item
tool-calling suite, under one fixed protocol for every row: temperature 0.6 / top_p 0.95 /
top_k 20, thinking forced on, a 16,384-token budget, concurrency 32, all rows back to back in one
sitting on one RTX PRO 6000 Blackwell.
Table with columns: build, size, overall, knowledge, call, abstain, runs k/a, finished, out/answer, tok/s @ 32, per-token latency| build | size | overall | knowledge | call | abstain | runs k/a | finished | out/answer | tok/s @ 32 | per-token latency |
|---|
| cyankiwi AWQ-INT4 | 72 G | 85.6 | 88.1 | 68.1 | 82.5 | 1/1 |
overall is one number over both suites: the 1,170 knowledge and 200 tool-calling items pooled
as 1,370, weighted 85.4% and 14.6% by item count. Repeat runs of one checkpoint moved it by
about a point, so treat a gap below 1.0 as a tie.
Table with columns: build, agentic, call, abstain| build | agentic | call | abstain |
|---|
| cyankiwi AWQ-INT4 | 71.0 | 68.1 | 82.5 |
| kkuspa NVFP4 | 67.5 | 65.6 | 75.0 |
| this repo | 67.9 | 64.6 | 81.0 |
|
call is accuracy on the 160 rows that require a call; abstain is the 40 whose correct action is
to call nothing. Benchmarks: BFCL v4 (30, including irrelevance), xLAM/APIGen (45), ToolACE
(45), Glaive v2 (40), nvidia When2Call (40, the abstention rows). Tool schemas go in the system prompt
and the model answers with a JSON array of calls, the same way for every model. An item counts only
if every expected call is present with acceptable arguments and no call is invented.
Abstention is the weak axis for every model we have measured (52–82%), so a build can look strong
on overall and still over-call.
The last row is a drift control: the same checkpoint re-measured after four hours of other work came
back within 0.5% on throughput and 0.2 points on the knowledge half. The agentic half was not
re-run on this model, which is why that row carries no overall.
Thinking is forced on for every row here, and it has to be. olka-fi's checkpoint ships a chat
template that defaults enable_thinking to false; left at its default it answers in 198 tokens and
posts a score that describes a different product, not a faster one. Every row in this table reasons
because the protocol makes it, not because the templates agree.
How to read this. The AWQ build is 1.3 points ahead on overall, at the edge of this suite's
±1-point spread, and it costs 8 GiB more memory, 6% less throughput at concurrency 32, and 30% more
tokens for every answer. This build takes the size, the aggregate throughput and the token bill. If a
point of accuracy is worth more to you than any of those, the AWQ build is the better choice; no
single build here wins every column.
Tool calling is the whole field's weak spot on this model, ours included. Every build lands
between 67.5 and 71.0 on the 200 tool-calling items, and every one of them declines in prose rather
than returning an empty call list, which is correct behaviour and scored as such here. If tool calling is
your workload, this is a model-level limit rather than a quantization one.
Comparable with our other models
Accuracy numbers move for reasons that have nothing to do with the model: a shorter token budget, a
different temperature, or whether the model was allowed to reason at all. So every number in this
table, on this card and on our other cards, comes from one fixed protocol.
The same 1,370 items: a 1,170-item knowledge suite (MMLU-Pro, ARC-Challenge, HellaSwag, WinoGrande,
CommonsenseQA, BoolQ, OpenBookQA, GSM8K, MATH-500) and a 200-item tool-calling suite (BFCL v4,
xLAM/APIGen, ToolACE, Glaive v2, nvidia When2Call). temperature 0.6, top_p 0.95, top_k 20,
thinking forced on, a 16,384-token budget, no reasoning parser, scoring the last ANSWER: in the
reply. Concurrency 32 on one RTX PRO 6000 Blackwell, each model's rows in one sitting. Auto-scored,
no LLM judge. Both halves are means of at least three runs per build.
Table with columns: model, shape, size, overall, knowledge, call, abstain, finished, out, tok/s @ 32| model | shape | size | overall | knowledge | call | abstain | finished | out | tok/s @ 32 |
|---|
| Laguna-XS-2.1 | 31 B MoE | 19.3 GiB | 81.7 | 83.8 | 68.4 | 73.5 | 98.9% |
Read overall with finished. overall scores an answer that overran the token budget as wrong,
but it cannot say whether the model needed the room or failed to stop; finished and out separate
those. A gap under 1.0 is a tie. The tok/s column comes from each model's own sitting and drifts
a few percent between sittings, so read it as a bracket.
call and abstain are the tool-calling suite's two halves, reported separately. call is
accuracy on the 160 items that require a tool call; abstain is the 40 whose correct action is to
call nothing. They used to be pooled into one agentic number, and the pooling misled: a model with
ordinary call accuracy and unusual abstention discipline outscored models that are better at
actually making calls. Weight them by your own workload's mix.
Table with columns: benchmark, Laguna-XS-2.1, Nemotron-3.5-Lightning-30B-A3B, Ornith-1.5-35B-A3B, Muse-Glimmer-30B, Qwen3.8-27B, Laguna-S-2.1, Qwen3.8-Flash-Next| benchmark | Laguna-XS-2.1 | Nemotron-3.5-Lightning-30B-A3B | Ornith-1.5-35B-A3B | Muse-Glimmer-30B | Qwen3.8-27B | Laguna-S-2.1 | Qwen3.8-Flash-Next |
|---|
| knowledge | | | | | | | |
| mmlu_pro | 79.0 | 82.0 |
What's quantized to what
Table with columns: tensors, format| tensors | format |
|---|
| routed expert projections | NVFP4 (group 16) |
| attention projections, shared-expert projections, the first dense layer | MXFP8 (group 32) |
embeddings, lm_head, router, attention gates, norms | BF16 |
Weights-only round-to-nearest, no calibration data required or embedded.
Revision note: if you pulled this repo before 19 Aug 2026
The first published revision had a packaging bug and you should re-pull. vLLM fuses each expert's
gate_proj and up_proj into one tensor, which carries exactly one NVFP4 global scale; that
revision wrote the two halves' scales independently, so 11,815 of 12,032 expert pairs disagreed,
the worst by a factor of 6.9, and one half of most experts was served mis-scaled. Every tensor was
correct on disk, so it loaded and mostly worked. The visible symptom was a failure to stop, leaving
21.5% of answers unfinished at a 16k budget.
The current revision fixes it, at the same size and the same format map:
Table with columns: acc, finished, out/answer | acc | finished | out/answer |
|---|
| first revision | 78.4 | 78.5% | 4,119 |
| current | 87.1 | 97.3% | 995 |
If you worked around it with repetition_penalty, you no longer need to.