Usage
vllm serve cloudnathan5/Qwen3.8-27B-NVFP4a16-GPTQ --max-model-len 32768 --max-num-seqs 512
--max-num-seqs matters on this architecture. 48 of the 64 layers use
linear attention, and vLLM allocates one Mamba-style cache block per decode
sequence. vLLM's default max_num_seqs=1024 can exceed the number of blocks
that fit, and startup then fails during CUDA graph capture with
max_num_seqs (1024) exceeds available Mamba cache blocks. Lower
--max-num-seqs (512 is a safe starting point) or raise
--gpu-memory-utilization. This is a property of the base model, not of
quantization.
from vllm import LLM, SamplingParams
llm = LLM(model="cloudnathan5/Qwen3.8-27B-NVFP4a16-GPTQ")
out = llm.generate(
["Explain 4-bit quantization in two sentences."],
SamplingParams(temperature=0.7, max_tokens=256),
)
print(out[0].outputs[0].text)
Single-user latency
Concurrency 1 (one request in flight — no batching), on a single NVIDIA RTX PRO 6000 Blackwell 96GB, vLLM 0.27.1. Prefix caching disabled and --ignore-eos set so prefill is never skipped and every generation is exactly 256 tokens. Median over 24 requests after 4 warmups, via vllm bench serve --max-concurrency 1.
Table with columns: input tokens, TTFT, inter-token latency, decode tok/s, BF16 base tok/s| input tokens | TTFT | inter-token latency | decode tok/s | BF16 base tok/s |
|---|
| 1024 | 191 ms | 20.0 ms | 50.1 (1.91x) | 26.2 |
| 4096 | 691 ms | 20.1 ms | 49.8 (1.91x) | 26.1 |
This is single-stream interactive performance, not batched throughput; under concurrency the ranking between variants differs.
Quantization details
Calibrated on 256 samples of
HuggingFaceH4/ultrachat_200k
at 4096 tokens, chat template applied.
The following modules are left in their original precision:
Table with columns: pattern, reason| pattern | reason |
|---|
lm_head, embed_tokens | quantizing these costs accuracy for no speed benefit |
visual.* | the vision tower is small and quantization-sensitive |
linear_attn.* | the gated-delta / linear-attention state paths are numerically fragile at 4 bits and are not GEMM-bound |
mlp.gate, shared_expert_gate | MoE routing weights |
mtp.* |
Reproduce with quantize.py:
python quantize.py --model-id Qwen/Qwen3.8-27B --method nvfp4a16-gptq
Evaluation — perplexity
Token-level perplexity on wikitext-2-raw-v1 (test), 48 non-overlapping 4096-token windows (196,560 tokens scored), measured through vLLM with identical settings for both rows.
Table with columns: model, perplexity, vs BF16| model | perplexity | vs BF16 |
|---|
| this checkpoint | 6.7129 | +2.37% |
| Qwen/Qwen3.8-27B (BF16) | 6.5574 | — |
This is token-level perplexity over a fixed window, which is not the same statistic as lm-eval's word_perplexity — compare it only against numbers produced the same way.
Caveats
- Quantization is lossy. Validate on your own workload before production use.
- The exclusion list above was derived from the architecture at release; if you
fine-tune or otherwise alter module naming, re-derive it.