Model Size Comparison
Table with columns: Variant, Approximate Size| Variant | Approximate Size |
|---|
| Qwen/Qwen3.8-27B (Full Precision BF16 / FP16) | ~54 GB |
| prithivMLmods/Qwen3.8-27B-FP8 (Compressed FP8) | ~36 GB |
Quantization Details
Quantization was performed using llmcompressor with the following recipe:
default_stage:
default_modifiers:
QuantizationModifier:
targets: [Linear]
ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*',
're:.*linear_attn.*']
scheme: FP8_DYNAMIC
bypass_divisibility_checks: false
requires_calibration_data: false
Linear layers are quantized to FP8 with dynamic per-tensor activation scaling, so no calibration dataset is required (requires_calibration_data: false). The lm_head, embedding table, any vision-tower (visual) components, and linear_attn layers are excluded from quantization and remain at full precision to preserve output-head fidelity and numerical stability.
Table | |
|---|
| Base model | Qwen/Qwen3.8-27B |
| Quantization scheme | FP8_DYNAMIC (Linear layers only) |
| Format | compressed-tensors |
| Calibration data required | No (dynamic activation scaling) |
| Excluded from quantization | lm_head, embed_tokens, visual (if present), linear_attn |
Use with vLLM
Qwen3.8-27B-FP8 is served through vLLM with native support for compressed-tensors FP8 checkpoints.
Requirements
torch >= 2.11.0
- A GPU with FP8 support recommended (Hopper or Blackwell class) for best throughput; also runs on Ampere with FP8 dequantized on the fly.
Serve
vllm serve prithivMLmods/Qwen3.8-27B-FP8 \
--max-model-len 32768
Client request
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
messages = [
{
"role": "user",
"content": "Explain how a transformer model processes text."
}
]
response = client.chat.completions.create(
model="prithivMLmods/Qwen3.8-27B-FP8",
messages=messages,
temperature=0.0,
max_tokens=512,
)
print(response.choices[0].message.content)
pip install transformers accelerate
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model = AutoModelForCausalLM.from_pretrained(
"prithivMLmods/Qwen3.8-27B-FP8",
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
"prithivMLmods/Qwen3.8-27B-FP8"
)
messages = [
{
"role": "user",
"content": "Explain how a transformer model processes text."
}
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt"
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=512
)
print(
tokenizer.decode(
outputs[0][inputs.shape[-1]:],
skip_special_tokens=True
)
)
Acknowledgments
- llmcompressor: Used to produce the FP8 dynamic quantization for this release.
- vLLM: Recommended inference engine with native compressed-tensors FP8 support.