What moved—and what did not
- K3.25 target experts: layers 3…44 start at K3, then exactly 9,072 of
36,288 projections move to K4. The promotion budget is
gate/up/down = 1,701 / 2,835 / 4,536 (the 3:5:8 quality allocation).
- Straight-K3 MTP: all 864 checkpoint-MTP projections in layer 45 stay K3.
It remains valid for diagnostics, but the serving recipe defaults to
DFlash2, so extra K4 bytes and search time are not spent there.
- 37,152 final quantized projections were produced in one
target-plus-MTP
model.quantize(...) call.
- Attention, dense MLPs, shared experts, routers, norms, embeddings, vision,
and all other tensors remain native precision.
- The release audit compared 1,618 native
tensors / 18.01 GiB byte-for-byte against source revision
f12e0fe1f6b2ea274c11a569582edfd99d993c5e.
- Packed checkpoint payload: 136.16 GiB across
18 safetensors shards.
Quantization used GPTQModel a053382584fa58cba7bf212ef1b829d08b29b2c0, EXL3 MCG adjacent tiers,
seed 787, sigma_reg=0.025, automatic output-scale selection, and a fixed
1,426-record calibration corpus (sha256:4e569625d97865777da92167b8fbf6fabb4ab7adf55baa5d676c81ce8dd95244).
Within every target layer, K4 promotion follows measured Hessian-weighted
relative error times natural gate-squared mass while the global 3:5:8 budget
stays exact. Natural GLM router traffic supplied expert Hessians; the committed
recovery contract covers any expert below the 1,024-route floor. Allocation
provenance and the validation report ship with the model; the full error ledger
is retained separately as internal quantization evidence.
Qualified two-GPU serving result
Release v0.7.0 of the matching B12x/vLLM recipe qualified the public target
revision with DFlash2 K5, FP8 MLA, TP2 + EP2 + DCP2, vision up to 16 images,
and a 1,048,576-token request limit on 2× RTX PRO 6000 Blackwell at a 400 W
power cap per GPU.
- Code-agent decode: 213 tok/s at C1 and 832 tok/s at C16.
- 128K cold prefill: 4,572 tok/s.
- Available KV cache: 2,758,919 tokens.
- Exact 1M multi-needle test: 6/6 needles recovered.
- Exact 69-case thinking tool-call suite at parallelism 8: 86/100
(55 pass, 8 partial, 6 fail), versus 88/100 for the matched uniform-K3
control.
The full prompts, outputs, scoring, performance curves, and machine-readable
receipts are in the recipe's
v0.7.0 benchmark report.
Serving
This is intended for the optimized two-GPU GLM-5.3 recipe at
tpurtell/glm-5.3-flash-ext3-4-bit-2x-rtx.
That runtime carries the GLM/EXL3/MTP and B12x ports; this card does not claim
stock-vLLM support.
Huge thanks to Z.ai for GLM-5.3 Flash, Brandon for the earlier K4 quant
and qualification work that helped inform this run, MiaAI-Lab for the nearby
dual-DGX-Spark reference, and the GPTQModel, ExLlamaV3, vLLM, and B12x
contributors.
Audit identity
- Source index:
sha256:e6007bd58fb7e07f9fe69544257ee2713f252ef5855bbf685b48c991d524ef0f
- One-shot plan:
sha256:cb8ee0b49c6a41bb31e9813214193fa9426c209dd0a0ca9f52ce8b5ab61465d8
- Validation:
sha256:30e5f8f5fc6c1d10cdd17749b40559a0dca89bec89b5affad219cab0f0260160