What is quantized
Table with columns: component, treatment| component | treatment |
|---|
| routed experts (layers 3–78, incl. MTP-78) | EXL3 trellis, per layer 256 experts K3 (avg 3.0 bpw), mcg codebook |
| dense MLP (layers 0–2), all attention, norms, embeddings, lm_head, mlp.gate, eh_proj | BF16, carried byte-exact |
| shared experts | BF16 in-checkpoint (online K6 at serve) |
- Data-free encode: identity Hessian (H=I), rotations + trellis search only —
no calibration capture anywhere. Deterministic: seeds derived from
(layer, expert, projection, rank), seed_base 20260711.
- K4 selection: per layer, the experts with highest relative round-trip MSE
under the K3 encode are re-encoded at K4 (downward remix: the 192 never-promoted K3 experts + the 64 banked donor-K3 payloads from the mix ledger — zero re-encode, byte-identical to a PROMOTE_K4=0 run).
- Per-expert tier map in
tier_bitmap.json; encode provenance in
config.json.hybrid_tr3_tail; file hashes in MANIFEST.sha256.
Serving
TR3 mixed-bit layout (carrier BF16 shards + per-layer trellis payloads) — serve
with an exllamav3-b12x/sparkinfer-lineage stack, TP4 (+ DCP4/MTP3), FP8 KV.
The mixed-K projection-tiers patch is REQUIRED; a stock loader that assumes a
uniform K per layer will produce fluent garbage. Not loadable by vanilla
exllamav3 model loading.
KLD vs BF16 teacher
Full-vocabulary, teacher-forced KL(teacher || student) against sealed BF16
GLM-5.3 logits — 4 held-out windows × 2,047 positions × 154,880 vocabulary,
fp32 log-softmax both sides. Measured independently on two different
4× RTX PRO 6000 (96GB) machines:
Table with columns: weight quant, KV mode, this work, CN3 (@dareposte), Δ| weight quant | KV mode | this work | CN3 (@dareposte) | Δ |
|---|
| 3.42 bpw | fp8 | 0.024105 | 0.023966 | −0.6% |
| 3.25 bpw | fp8 | 0.026103 | 0.026776 | +2.6% |
| 3.25 bpw | nvfp4 |
Readings: the 3.25↔3.42 weight step changes KLD by only ~0.002; the
fp8→nvfp4 KV step costs ~7× more (~0.014) — cache format matters more than
the extra 0.17 bpw. Under nvfp4 KV the weight-quant difference washes out
entirely. The window SD (~0.02–0.03) is corpus heterogeneity — one
citation-dense legal window is uniformly hardest; dialogue, explanatory
prose, and reasoning-trace registers measure near-transparent.
Table with columns: config, source, w0000, w0001, w0002, w0003| config | source | w0000 | w0001 | w0002 | w0003 |
|---|
| 3.42 fp8 | this work | 0.0141 | 0.0537 | 0.0137 | 0.0148 |
| 3.42 fp8 | CN3 | 0.0148 | 0.0542 | 0.0135 | 0.0134 |
| 3.25 fp8 | this work | 0.0188 |
Method (reproducible)
- Teacher: brandonmusic/GLM-5.3-BF16-full-logits,
reference-full-panel confirmation lane (held out from every
calibration fit), revision 427368f1.
- Student: this checkpoint, loaded by the digest-pinned r17 serving
image (
sha256:c5e96c5b…) — the real trellis kernels and online-K6 path,
offline vllm.LLM, TP4, one teacher-forced prefill per window.
- Full runbook + runner:
kld/ in this repo
(KLD-REPRODUCTION.md, prefill_kld_53.py, fetch-teacher.sh).
- Independent reproduction bundle (receipts, unedited logs, checksums,
pinned revisions): .
Capacity caveat (via the CN3 report): KV-pool token figures printed by
KLD-profile boots (TP4/DCP1, 4,096-token envelope) are logical pool
values for that profile only — they are not maximum context length or
production serving capacity.
Credits
- brandonmusic — thank you for the
GLM-5.3-BF16-full-logits
teacher captures that make this measurement possible without a 1.5TB BF16
forward, and for the TR3 quantization references this release follows: the
GLM-5.2-EXL3-TR3v4-3.5bpw-MTP78
method card, runbook, and the r10 reproducibility bundle whose encoder
lineage (
encode_tr3_v31.py) this checkpoint was produced with.
- dareposte — thank you for the
independent CN3 reproduction: all six weight/KV configurations on separate
hardware, within ±5% of our means, published here with receipts under
kld/cn3/.
- willfalco/GLM-5.2-EXL3-TR3-3.25bpw
— the 3.25 tier recipe and mixed-K checkpoint-format lineage.
Status
Context-maximal artifact of the release trio (3.42 quality / 3.25 balanced / 3.0 breadth): ~24 GiB lighter than 3.25 => roughly +530k KV tokens at fp8. Structurally gated at assembly (58,368 promoted slots shape-verified K3). KLD measurement to follow — expected ~0.030 fp8 by ladder extrapolation; the reproduction kit in kld/ scores it unmodified.
Quantized with encode_tr3_53.py (exllamav3 v0.0.43 vendored math, MIT) on
4x RTX PRO 6000 Blackwell. Encode + tooling notes ship in the repo.