Identical layout to the official zai-org/GLM-5.3-Flash FP8 release:
quant_method: fp8, fmt: e4m3, activation_scheme: dynamic, weight_block_size: [128, 128]
- Exactly the same set of tensors is quantized as in the official checkpoint (routed + shared experts, dense MLPs, MLA projections); everything in
modules_to_not_convert (embeddings, lm_head, routers, norms, linear-attention/KDA, hyper-connections, vision tower) is left in BF16/F32 as in the source.
- Each FP8 weight has a
weight_scale_inv (F32, one scale per 128×128 block, scale = amax / 448). Round-to-nearest, no calibration data.
- Tensor names and the
quantization_config are copied from the official release, so any engine that serves zai-org/GLM-5.3-Flash should load this checkpoint the same way.
Serving
Use the same setup as for zai-org/GLM-5.3-Flash, e.g.
vllm serve pqhaz/apex-flash-1-abliterated-FP8 --tensor-parallel-size 4
python -m sglang.launch_server --model-path pqhaz/apex-flash-1-abliterated-FP8 --tp 4
Notes
- From the upstream card: this abliterated variant has not undergone a separate evaluation. Results reported for apex-flash-1 apply to the standard checkpoint only. It is intended for authorized security research.
- This quantization has not been separately benchmarked either.
- Not affiliated with Cantina Security or Z.AI.
License
MIT, inherited from the base model. Copyright (c) 2026 Z.AI Co., Ltd — see LICENSE.