Architecture
graph TD
Base["Kwaipilot/KAT-Coder-V2.5-Dev<br/>Qwen3.5 MoE - 256 experts - ~69 GB bf16"]
subgraph Build ["Build - RTX 5070 Ti, SM120"]
REAP["REAP expert prune 50% 256 -> 128 experts + router-renorm fix"]
Strip["strip vision tower + 333 untrained tensors"]
Quant["NVFP4A16 quantize (weight-only, data-free, 82 s)"]
end
subgraph HF ["Published formats"]
A16["REAP-50-NVFP4A16 - 12.45 GiB (default, vLLM)"]
W4A4["REAP-50-NVFP4-W4A4 (native FP4)"]
GPTQ["REAP-50-NVFP4A16-GPTQ (null result, kept for transparency)"]
GGUF["REAP-50-GGUF (Q4_K_M / Q5_K_M / Q6_K / Q8_0)"]
BF16["REAP-50-bf16 (pruned source)"]
end
Bench["A16 single draws - HumanEval+ 89.0% - MBPP+ 90.5% - SWE-bench Verified 52.0%"]
Base --> REAP --> Strip --> Quant --> A16
Strip --> BF16
BF16 -. re-quant .-> W4A4
BF16 -. re-quant .-> GPTQ
BF16 -. convert .-> GGUF
A16 --> Bench
Why two builds
Weight-only quantization (NVFP4A16) dequantizes on the fly and computes in
bf16 — there is no hardware path for a weight-only scheme to reach a GPU's
native FP4×FP4 tensor-core instructions, on any GPU. Quantizing activations
too (this build) does reach those native kernels
(FlashInferCutlassNvFp4LinearKernel for dense layers, FLASHINFER_CUTLASS
for MoE experts, confirmed via vLLM's own kernel-selection logging). The
question this build answers: does native FP4×FP4 compute actually beat a
dequant-then-bf16-compute fallback on consumer Blackwell (SM120, e.g. RTX
5070 Ti/5080/5090) for a real agentic-coding decode workload?
Measurements
5 interleaved process invocations per arm (warmup discarded), median + range
reported, under both isolated eager-mode execution (cleanly isolates kernel
dispatch, but understates real throughput by ~7x on this hardware) and the
same PIECEWISE CUDA-graph configuration the NVFP4A16 model serves with in
production:
Table with columns: HumanEval, HumanEval+, MBPP+, decode, eager, decode, PIECEWISE (production) | HumanEval | HumanEval+ | MBPP+ | decode, eager | decode, PIECEWISE (production) |
|---|
| NVFP4A16 (sibling release) — single draw | 95.7% | 90.9% | 89.9% | 18.8 tok/s | 142.5 tok/s |
| W4A4 (this build) — median of 5 draws | 93.90% | 89.63% | 89.68% | 14.5 tok/s | 119.2 tok/s |
| W4A4 observed range over 5 draws | 92.07–96.34% | 87.20–90.85% | 88.62–90.48% | — | — |
On this hardware, for this single-stream (batch=1) workload: NVFP4A16 is
faster (0.84x for W4A4 under the production PIECEWISE configuration, 0.77x
in isolated eager mode). Both numbers are fast in absolute terms — 119
tok/s is roughly 17-20x typical human reading speed and comfortably
interactive for coding use; the gap is relative to the sibling build, not a
usability threshold.
Accuracy: we make no claim in either direction between the two builds.
These scores are not deterministic, and an earlier version of this card treated
them as if they were. Greedy decoding (do_sample: false, --seed 1234) fixes
token selection but not batch composition: vLLM's continuous batching groups
requests differently on each run and FP4 reductions are not order-invariant, so
logits differ in their last bits and the argmax occasionally tips. Re-running
this checkpoint five times with byte-identical arguments:
Table with columns: task, median, range over 5 draws, spread| task | median | range over 5 draws | spread |
|---|
| HumanEval | 93.90% | 92.07 – 96.34% | 4.27 pp |
| HumanEval+ | 89.63% | 87.20 – 90.85% | 3.66 pp |
| MBPP+ | 89.68% | 88.62 – 90.48% | 1.85 pp |
Every difference between the A16 and W4A4 rows above is smaller than the
run-to-run spread of a single row. The previous card read those differences as
a pattern — "lower on HumanEval and HumanEval+, slightly higher on MBPP+" — and
that was not supportable. It has been withdrawn rather than restated more
softly.
Two caveats stated plainly. The A16 row is a single draw from 2026-08-20 and has
not been re-run five times; treat it as one sample from a distribution of
comparable width. And the MBPP+ figure this card previously published, 91.01%
(344/378), lies above all five draws measured on 2026-09-05 (max 342/378) — it
is not reproducible from this checkpoint and has been replaced by the median.
The throughput figures are unaffected: those were measured with 5 interleaved
reps per arm with median and range reported, and the 0.84x gap sits far outside
their spread.
Reproduce any of this with scripts/eval/eval_repeat.sh in the project repo,
which runs a task N times and reports median, range and a Wilson interval.
For reference, QSpec-era literature (INT4-generation quantization) reported
W4A4 losing up to 38.73% on HumanEval; that collapse did not reproduce here,
consistent with NVFP4's per-16-block scaling and FP8 scale factors being a
better-conditioned format than INT4-era quantization.
Same checkpoint size either way (12.4532 GiB here vs 12.4512 GiB for A16).
Which build to use
- NVFP4A16 (the sibling release) is faster in our measurements and is
the one we default to for our own agentic-coding pipeline.
- This W4A4 build is the one to reach for if you specifically want the
native FP4×FP4 tensor-core execution path (e.g. building on top of
activation quantization, or targeting a serving stack where that path
matters more than it did for us), or if you want to independently verify
or extend the comparison above.
We only measured single-stream (batch=1) decode, matching our own agentic
use case. We have not measured batch sizes above 1, where the native
kernel's throughput/latency characteristics may differ from the dequant
path's — if you test that, we'd like to hear what you find.
Mechanism notes
Two findings surfaced while measuring this, independent of the headline
numbers:
- At identical
gpu_memory_utilization, the NVFP4A16/Marlin arm needed a
higher utilization setting to reliably allocate KV cache than this W4A4
build did at the same checkpoint size — Marlin's dequantize-on-the-fly
path appears to need more non-weight runtime workspace than the native
FP4 kernel path does.
- W4A4 gains proportionally more from CUDA graphs than NVFP4A16 does
(eager→PIECEWISE: NVFP4A16 7.6x, W4A4 8.2x), which narrows the gap
between them (0.77x → 0.84x) without closing it.
Usage
Loads and serves like the NVFP4A16 sibling (same architecture,
transformers/vllm requirements, SM120/compute-capability-12.0
requirement, no CPU offload needed at 12.45 GiB). See the sibling release's
model card for full serving instructions, environment requirements, and the
REAP-pruning background — this card documents only what differs about this
build.
from vllm import LLM, SamplingParams
llm = LLM(
model="Ttimms/KAT-Coder-V2.5-Dev-REAP-50-NVFP4-W4A4",
dtype="bfloat16",
language_model_only=True,
trust_remote_code=True,
)
License
Apache 2.0, inherited from the base model Kwaipilot/KAT-Coder-V2.5-Dev and
its upstream lineage.