📖 Cookbook — setup, patches, methodology, evidence
Conversion contract
- 4-bit, symmetric (
sym: true), group size 64, desc_act: false
- Signed range
[-8, 7], stored nibble signed+8, low-first
qweight [K/8, N] uint32; scales BF16 [K/G, N]; qzeros zero-filled
- Excluded (kept BF16): embeddings,
lm_head, 1D norms
- Expert geometry preserved: 128 routed + 1 shared, non-gated, heterogeneous
intermediate sizes (1856 / 3712)
- Per-matrix error stats:
conversion-manifest.json
Companion DFlash draft
For the measured speculative path, pair this target with
SergiioB/Nemotron-3.5-Lightning-30B-A3B-DFlash-BF16
(local NVFP4→BF16 reconstruction of NVIDIA's official DFlash draft).
Measured on Intel Arc Pro B70
Real measurements from the publisher's card: one Arc Pro B70 32 GB,
150 W configured cap, vLLM 0.26.1rc1.dev668+g3ee2df303 XPU, C1, prefix
cache off, client-side timing, n=5 medians. Commands and raw logs are in
the cookbook (links above).
Table with columns: Mode, Cell, Metric, median, notes| Mode | Cell | Metric | median | notes |
|---|
| no-spec + XPU graphs | p512/g128 | C1 client post-first | 93.00 t/s | range 92.96–93.03; eager was 21.8 |
| no-spec + XPU graphs | p8192/g128 | C1 client post-first | 87.25 t/s | range 87.22–87.31 |
| DFlash n_spec=7 | p2048/g128 | C1 client post-first | 186.61 t/s |
An earlier ~10.3k figure was a no-spec n=3 TTFT-derived cold input rate
on a decode cell (p8192/g128, median TTFT 0.7916 s → 10,349). It is not
the isolated DFlash n=5 input number and it is not engine prefill.
Caveats
- Speed is not task-quality or official-quant parity.
- Native MTP on this stack historically accepts 0% drafts. Use DFlash.
- Temperature-0 replay on the compiled/graph path has an upstream XPU
caveat in general; the isolated DFlash n=5 smoke matched.
License
OpenMDW-1.1, same as the NVIDIA source (LICENSE included).