GLM-5.3 NVFP4A16 with MTP — PR3118 validation
Full pretrained checkpoint converted from zai-org/GLM-5.3 using llm-compressor PR3118, 572c6390.
Checkpoint structure and MTP tensor checks passed. Runtime validation and acceptance-rate measurement are pending. No quality or performance benchmark is claimed.
The backbone MLP gate/up/down projections and supported MTP attention and MLP projections use data-free NVFP4A16: weight-only FP4 with 16-bit activations. Other projections remain BF16. Native FP8 source weights are dequantized before conversion. The backbone uses model_free_ptq with two GPUs, followed by PR3118’s prepare_mtp_save / save_mtp_tensors integration for MTP. The model reference contains only meta tensors; this run does not test a full-model oneshot load. This run replaces MFPTQ’s legacy per-shard fusion diagnostic with complete-index partner and module-identity checks; each shard worker also explicitly sets its assigned CUDA device for Triton packing. Quantization math and dependency resolution are unchanged. See config.json and pr3118-validation.json for the configuration and checks.
The output generation defaults enable sampling to make the source top_p=0.95 setting valid under Transformers 5.17.
The source model's licensing applies; consult its model card and included license files.
This artifact is intended for PR validation. It can be downloaded later for runtime testing.