Model Size Comparison
Table with columns: Variant, Approximate Size| Variant | Approximate Size |
|---|
| Qwen/Qwen3.8-27B (Full Precision BF16 / FP16) | ~54 GB |
| prithivMLmods/Qwen3.8-27B-FP8 (Compressed FP8) | ~36 GB |
Quantization Details
Quantization was performed using llmcompressor with the following recipe:
default_stage: default_modifiers: QuantizationModifier: targets: [Linear] ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*', 're:.*linear_attn.*'] scheme: FP8_DYNAMIC bypass_divisibility_checks: false requires_calibration_data: false
Linear layers are quantized to FP8 with dynamic per-tensor activation scaling, so no calibration dataset is required (requires_calibration_data: false). The lm_head, embedding table, any vision-tower (visual) components, and linear_attn layers are excluded from quantization and remain at full precision to preserve output-head fidelity and numerical stability.
Table | |
|---|
| Base model | Qwen/Qwen3.8-27B |
| Quantization scheme | FP8_DYNAMIC (Linear layers only) |
| Format | compressed-tensors |
| Calibration data required | No (dynamic activation scaling) |
| Excluded from quantization | lm_head, embed_tokens, visual (if present), linear_attn |
Use with vLLM
Qwen3.8-27B-FP8 is served through vLLM with native support for compressed-tensors FP8 checkpoints.
Requirements
torch >= 2.11.0
- A GPU with FP8 support recommended (Hopper or Blackwell class) for best throughput; also runs on Ampere with FP8 dequantized on the fly.
Serve
vllm serve prithivMLmods/Qwen3.8-27B-FP8 \ --max-model-len 32768
Client request
from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") messages = [ { "role": "user", "content": "Explain how a transformer model processes text." }] response = client.chat.completions.create( model="prithivMLmods/Qwen3.8-27B-FP8", messages=messages, temperature=0.0, max_tokens=512,) print(response.choices[0].message.content)
pip install transformers accelerate
from transformers import AutoTokenizer, AutoModelForCausalLMimport torch model = AutoModelForCausalLM.from_pretrained( "prithivMLmods/Qwen3.8-27B-FP8", torch_dtype="auto", device_map="auto") tokenizer = AutoTokenizer.from_pretrained( "prithivMLmods/Qwen3.8-27B-FP8") messages = [ { "role": "user", "content": "Explain how a transformer model processes text." }] inputs = tokenizer.apply_chat_template( messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to(model.device) outputs = model.generate( inputs, max_new_tokens=512) print( tokenizer.decode( outputs[0][inputs.shape[-1]:], skip_special_tokens=True ))
Acknowledgments
- llmcompressor: Used to produce the FP8 dynamic quantization for this release.
- vLLM: Recommended inference engine with native compressed-tensors FP8 support.