Quantization
Table with columns: Component, Method, Weight format, Details| Component | Method | Weight format | Details |
|---|
| Decoder Linear layers | AutoRound | symmetric INT4, G256 | lm_head excluded; 200 iterations; batch size 1 |
lm_head | GPTQ | symmetric INT8, per-channel | static act-order |
| Mixed-precision exceptions | — | source dtype | non-Linear parameters remain at source precision |
Calibration used HuggingFaceH4/ultrachat_200k, split train_sft[:512].
Conversations were rendered with the source chat template and packed into 512
sequences of 1,024 tokens without shuffling.
Evaluation
Full wikitext-2-raw-v1 evaluation used the wikitext lm-eval task with no
example limit and matched source/quantized settings on 2026-07-31.
Table with columns: Checkpoint, Word perplexity, Status| Checkpoint | Word perplexity | Status |
|---|
microsoft/Phi-4-mini-instruct | 11.713516 | Full run |
| This quantized checkpoint | 12.560528 | Full run |
| Absolute degradation | 0.847012 | Lower is better |
| Relative degradation | 7.231% | 100 * (quantized / source - 1) |
Finite-scale validation and a Transformers chat-generation smoke test passed.
Reproduction
This directory includes the model-specific quantize.py, recipe.yaml, and
versions.txt.
python quantize.py \
--model-path /path/to/models--microsoft--Phi-4-mini-instruct \
--output-dir /path/to/Phi-4-mini-instruct-Autoround-Safetensors
Environment
Exact Python, CUDA, Torch, Transformers, llmcompressor, AutoRound, and
compressed-tensors versions are recorded in versions.txt.
Deployment
This is the pre-LLiMa quantized Hugging Face artifact. Compile it separately
for the target Sima.ai platform and keep compiler output separate.
Limitations
Quantization quality can vary by language, domain, prompt format, context
length, and deployment runtime. Validate the intended workload independently.