Calibration and quantization
The quantization used 1,426 prompts totaling 1,081,027 tokens. The corpus
emphasizes English and Chinese and combines code, multilingual general
knowledge, worked math, reasoning/termination, and structured tool-call
examples. Code spans Python, C++, CUDA, C, Rust, and a broader mix of systems
and application languages.
This is a direct single-pass mixed quantization, not a blend of separately
generated K2 and K3 models. Every projection is encoded at K2, scored by its
Hessian-weighted relative error times natural gate-squared mass, and the
highest-scoring projections are immediately re-encoded at K3 within a fixed
gate:up:down = 3:5:8 budget. The mixed block is replayed before advancing,
so every later layer is calibrated against the real mixed prefix.
Natural top-6 router selections supplied expert activations. Experts below
the 1,024-row target received deterministic candidates from router ranks
7–12, with any remaining deficit represented by an explicitly counted
isotropic Hessian residual. The three dSpark blocks were replayed from a
327,680-anchor stratified sample with all five proposal rows issued jointly.
All 43 target blocks and all three dSpark blocks passed their independent
projection audits. The final checkpoint contains 35,328 routed projections:
3,302 of 33,024 target projections and 461 of 2,304 dSpark projections use K3.
DS4RT qualification
These results use one RTX 6000 Blackwell coordinator plus four DGX Sparks with
TP4 experts, the balanced FP8-KV profile, adaptive dSpark, greedy decoding,
and the same current DS4RT build for both artifacts.
Table with columns: Measurement, This model, K2 calibrated v1 control| Measurement | This model | K2 calibrated v1 control |
|---|