Why another NVFP4 reranker
The point of a 4-bit reranker is to leave VRAM for something else — a main model's KV cache,
or a co-located embedding model. Existing NVFP4 checkpoints of this model spend ~0.72 GiB on a
tensor that is never used at inference. Verified by reading the safetensors headers directly:
Table with columns: checkpoint, tensor bytes, avoidable overhead| checkpoint | tensor bytes | avoidable overhead |
|---|
Forturne/Qwen3-Reranker-4B-NVFP4 | 3.350 GiB | stores lm_head.weight (bf16, 740.6 MiB) and model.embed_tokens.weight, although tie_word_embeddings: true |
throwerdopey0/Qwen3-Reranker-4B-seq-cls-NVFP4 | 3.351 GiB | stores model.embed_tokens.weight upcast to F32 (1481.1 MiB instead of 740.6 MiB in bf16) |
| this checkpoint | 2.627 GiB | — embedding table stored once, in bf16 |
That is ~0.72 GiB back. In a deployment where the freed memory becomes KV cache, at a typical
~54k tokens/GiB for a 27B-class model that is roughly 39 000 extra context tokens.
Nothing is wrong with those checkpoints numerically — this is purely about not shipping a
redundant copy of a 151669 × 2560 table.
What is quantized — and what deliberately is not
targets=["Linear"], ignore=["lm_head"]. All 36 decoder layers' q/k/v/o_proj and
gate/up/down_proj are FP4, weights and input activations alike. Embeddings, norms and the LM
head are untouched.
lm_head must not be quantized. vLLM converts this model to
Qwen3ForSequenceClassification and builds the ranking head out of two rows of lm_head —
the "no" and "yes" token embeddings (classifier_from_token). It computes
score_weight = lm_head[yes] − lm_head[no], loads that into a 1-row score layer, and then
does del lm_head. Quantizing lm_head therefore injects error directly into the only two
rows that produce the score. Because tie_word_embeddings: true, lm_head is
embed_tokens, which is an nn.Embedding and so is already outside targets=["Linear"] —
but the ignore entry keeps that explicit and survives future recipe edits.
The same choice is made by the reference FP8 checkpoint
DCC-BS/Qwen3-Reranker-4B-FP8-Dynamic.
Serving with vLLM
vllm serve KoliaNik/Qwen3-Reranker-4B-NVFP4 \
--hf-overrides '{"architectures":["Qwen3ForSequenceClassification"],"classifier_from_token":["no","yes"],"is_original_qwen3_reranker":true}'
vLLM does not apply the reranker prompt template for you — this model has no
get_score_template, so for "LLM as reranker" the score endpoint simply concatenates
text_1 + text_2 and tokenizes the result. Build the canonical prompt yourself and split it
across the two fields:
import requests
PREFIX = ('<|im_start|>system\n'
'Judge whether the Document meets the requirements based on the Query and the '
'Instruct provided. Note that the answer can only be "yes" or "no".'
'<|im_end|>\n<|im_start|>user\n')
SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
INSTRUCT = "Given a web search query, retrieve relevant passages that answer the query"
def text_1(query, instruct=INSTRUCT):
return PREFIX + f"<Instruct>: {instruct}\n<Query>: {query}\n<Document>: "
def text_2(doc):
return doc + SUFFIX
r = requests.post("http://127.0.0.1:8000/v1/score", json={
"model": "KoliaNik/Qwen3-Reranker-4B-NVFP4",
"text_1": text_1("What is the capital of France?"),
"text_2": [text_2("Paris is the capital and largest city of France."),
text_2("The blue whale is the largest marine mammal.")],
})
print([d["score"] for d in r.json()["data"]])
The separators are single \n, taken from the model's own chat_template.jinja. (The base
model card's vLLM snippet uses \n\n; the template does not. Pick one and keep it consistent
between indexing and evaluation.) /v1/rerank works the same way. The returned score is
sigmoid(logit_yes − logit_no), i.e. P("yes").
The template trap — read this before you benchmark
--hf-overrides builds the score head, but it does not make vLLM apply the reranker
prompt template. If neither side supplies it, vLLM receives a bare query glued to a bare
document. The engine still starts, still registers /v1/score and /v1/rerank, still returns
plausible scores in 0–1 — and silently ranks nonsense.
Measured on the bf16 base model with no template, for the query "What is the capital of
France?":
Table with columns: document, score, rank| document | score | rank |
|---|
| Berlin is the capital of Germany. | 0.800 | 1 |
| Paris is the capital and largest city of France. | 0.754 | 2 |
| The cat sleeps on the windowsill. | 0.383 | 3 |
The model separates on-topic from off-topic but cannot tell which capital was asked about,
because the query never lands in its slot. With the template applied, the same three documents
score 0.991 / 0.000007 / 0.0000004 and order correctly.
The tell is the token count: ~25 prompt tokens per pair without the template, 80–100 with
it. Two ways to supply it correctly:
- Server-side — pass
--chat-template <vllm>/examples/pooling/score/template/qwen3_reranker.jinja alongside
--hf-overrides (the vLLM example also passes --runner pooling). Then send plain queries
and documents; /v1/score and /v1/rerank additionally accept an instruction field.
- Client-side — build the canonical prompt yourself and split it so that
text_1 + text_2 reconstructs it exactly, as in the snippet above.
Either way, verify it: POST /tokenize on the concatenated prompt must return exactly the
usage.total_tokens that /v1/score reports. And after any config change make the smoke test
"does a known-relevant document outrank a known-irrelevant one", not "does the endpoint
answer 200" — the failure above passes the second check.
Every number in this card was produced with the client-side form and verified against
/tokenize, so none of them are affected by this trap.
Measurements
All numbers below were measured directly, on 2× RTX 5090 (SM120) with vLLM 0.28.0,
--tensor-parallel-size 2 --enforce-eager --max-model-len 768 --kv-cache-memory-bytes 134217728 --max-num-batched-tokens 768
. The GPUs were shared with
another resident workload, so these are
not peak-throughput figures — but every variant
ran under an identical configuration, so they are comparable to each other.
Size and VRAM
Table with columns: variant, file (tensors), weights in VRAM, per-card occupancy at TP=2| variant | file (tensors) | weights in VRAM | per-card occupancy at TP=2 |
|---|
| bf16 base | 7.491 GiB | 7.48 GiB | — |
| FP8-Dynamic (reference) | 4.833 GiB | 4.22 GiB | 3.35 GiB |
| NVFP4A16 (sibling) | 2.627 GiB | 2.70 GiB | 2.59 GiB |
| NVFP4 W4A4 (this) | 2.627 GiB | 2.68 GiB |
The FP8 file is 4.833 GiB but only 4.22 GiB reaches VRAM — vLLM drops lm_head after
building the score head. Per-card occupancy includes the CUDA context, KV cache and
activations. This checkpoint needs ~0.23 GiB/card more than A16 at runtime for the FlashInfer
workspace, despite identical weights on disk.
Composition of this checkpoint: 1.692 GiB packed FP4 (uint8) + 0.211 GiB FP8 group scales
- 0.724 GiB bf16 (embeddings and norms); the per-layer
input_global_scale tensors are
F32 scalars and round to zero.
Single-GPU gotcha. At TP=1 vLLM allocates a full ParallelLMHead
(151669 × 2560 bf16 = 742 MiB) before tying it to embed_tokens and deleting it. Load
peak therefore exceeds steady state by that much; at TP=2 it is split across cards. Budget
~0.75 GiB of transient headroom when starting this model on a single GPU.
Quality
Public reranking benchmarks, 150 sampled queries each, documents truncated to 800 characters.
Δ is the mean per-query difference against bf16 with its standard error; Spearman/Pearson and
mean |Δscore| are computed over all pairs.
MTEB RuBQReranking (Russian, 3535 pairs):
Table with columns: variant, nDCG@10, MRR, MAP, Δ nDCG@10 vs bf16, Spearman, mean |Δscore|| variant | nDCG@10 | MRR | MAP | Δ nDCG@10 vs bf16 | Spearman | mean |Δscore| |
|---|
| bf16 | 0.8773 | 0.8585 | 0.8165 | — | — | — |
| FP8-Dynamic | 0.8741 | 0.8546 | 0.8118 | −0.0032 ± 0.0050 | 0.9973 | 0.0089 |
MTEB SciDocs-reranking (English, 3600 pairs):
Table with columns: variant, nDCG@10, MRR, MAP, Δ nDCG@10 vs bf16, Spearman, mean |Δscore|| variant | nDCG@10 | MRR | MAP | Δ nDCG@10 vs bf16 | Spearman | mean |Δscore| |
|---|
| bf16 | 0.9362 | 0.9811 | 0.9056 | — | — | — |
| FP8-Dynamic | 0.9344 | 0.9813 | 0.9029 | −0.0019 ± 0.0017 | 0.9982 | 0.0053 |
For scale: scoring the same checkpoint twice gives Pearson 0.99975 and mean |Δscore| 0.0022
(vLLM batching is not bit-deterministic). Read honestly: on the Russian set this checkpoint is
the closest of the three to bf16, while on the English set it is the furthest and its
−0.0079 ± 0.0034 is the one deviation in these tables that clears its own error bar.
Independent in-house cross-check
Both quantizations were also run against a private Russian-language retrieval benchmark that
is not public and therefore not reproducible from this repository: 100 queries, top-50
candidates from a BM25 first stage, with a relevant document present among the candidates in
89% of cases.
Table with columns: variant, nDCG@10, R@1, R@5, MRR@10, s/query| variant | nDCG@10 | R@1 | R@5 | MRR@10 | s/query |
|---|
| first stage, no reranking | 0.643 | 0.45 | 0.79 | 0.587 | — |
| bf16 base | 0.758 | 0.570 | 0.89 | 0.713 | 0.82 |
| FP8-Dynamic | 0.767 | 0.590 |
On that set this checkpoint matched A16 on nDCG@10, took the best R@1 of everything measured,
and ran twice as fast per query as bf16 and A16 — the same FP4-activation speedup the
public throughput table shows, on a different configuration. Neither quantization lost anything
against bf16; both landed slightly ahead, within noise at 100 queries. Note the ceiling:
R@5 = R@10 = 0.89 for every Qwen variant, which is exactly the share of relevant documents the
first stage retrieved at all, so the reranker is not what limits the pipeline there.
Treat this as a directional cross-check, not as headline numbers: the corpus is private, the
sample is small, and the timings come from a different serving configuration than the table
below.
Throughput
RuBQReranking, 3535 pairs, warm run, 4 concurrent requests, identical engine settings:
Table with columns: variant, pairs/s, vLLM kernel on SM120| variant | pairs/s | vLLM kernel on SM120 |
|---|
| FP8-Dynamic | 103.5 | CutlassFP8ScaledMMLinearKernel |
| NVFP4A16 | 75.9 | MarlinNvFp4LinearKernel (weight-only) |
| NVFP4 W4A4 | 114.8 | FlashInferCutlassNvFp4LinearKernel |
This is the fastest of everything measured, and the only 4-bit variant that does not regress
against FP8. Worth stating plainly, because the opposite is easy to assume: a weight-only
NVFP4 checkpoint on this card is forced onto Marlin and logs "Your GPU does not have native
support for FP4 computation". That warning does not apply here — with quantized activations
the normal kernel-selection path runs and finds a working FlashInfer CUTLASS FP4 kernel on
SM120.
Which variant should I use?
Both this checkpoint and the A16 sibling occupy the same 2.627 GiB on disk, so the memory
argument does not separate them. What separates them is the kernel vLLM picks:
- W4A4 (this one) —
FlashInferCutlassNvFp4LinearKernel, fastest measured, no speed
regression against FP8. Costs ~0.23 GiB/card of runtime workspace and has the lowest score
fidelity of the three, with a small but real ranking loss on the English set.
- A16 — forced onto Marlin (weight-only), ~27% slower than FP8, but better fidelity,
smaller runtime footprint, data-free and reproducible in 24 seconds on a CPU.
Pick W4A4 if reranker throughput matters. Pick A16 if the reranker is a low-QPS
auxiliary component where the smallest runtime footprint and the most conservative numerics
win.
On the private domain set above the two were level on nDCG@10 while this checkpoint ran twice
as fast per query, so the English-set gap did not transfer there. If you have a domain
benchmark, run both — that measurement will decide this better than either table here.
How it was made
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
import torch
MODEL = "Qwen/Qwen3-Reranker-4B"
model = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16, device_map="cpu")
tok = AutoTokenizer.from_pretrained(MODEL)
oneshot(
model=model,
recipe=QuantizationModifier(targets="Linear", scheme="NVFP4", ignore=["lm_head"]),
output_dir="Qwen3-Reranker-4B-NVFP4",
processor=tok,
dataset=calibration_dataset,
num_calibration_samples=512,
max_seq_length=512,
pipeline="sequential",
)
In compressed-tensors, scheme="NVFP4" is W4A4: its input_activations are 4-bit with a
static_minmax observer, so calibration data is required. The data-free weight-only
scheme is NVFP4A16, whose input_activations is None. Getting these two confused is easy
and the failure is silent.
Calibration set: 512 prompts built in the exact serving format above from
mteb/RuBQReranking and mteb/scidocs-reranking (400 queries from each), truncated to 512
tokens. Queries were filtered to be disjoint from the evaluation sample, so the numbers
above are not calibrated-on-test. Everything is public data.
Ran on CPU (24 cores) in 104 minutes. llm-compressor logs
Re-tied input/output embeddings; saving a single shared table
— that is the line that keeps the checkpoint at 2.627 GiB
instead of 3.35 GiB.
Evaluation methodology
Datasets: mteb/RuBQReranking and
mteb/scidocs-reranking, test
splits, 150 queries sampled with seed 20260907, up to 24 candidates per query, documents
truncated to 800 characters so that every variant sees byte-identical prompts.
Scores are P("yes") from the 2-way softmax over the "no"/"yes" logits at the final
position — identical by construction to what vLLM's from_2_way_softmax pooler computes. The
harness was cross-checked against an independent transformers implementation on CPU in
float32: Pearson 0.999967, mean |Δscore| 0.0007.
The bf16 baseline was served through the same vLLM engine (with --cpu-offload-gb 2, since
7.48 GiB of weights did not fit in the available VRAM), so all four rows in the tables come
from one engine and one prompt builder.
Limitations
- Throughput figures come from GPUs shared with another workload, with a deliberately tiny KV
cache and
--enforce-eager. Treat them as a relative ranking, not as absolute performance.
- Measured only on SM120 (RTX 5090). On SM100 (B200) kernel selection differs, and so will the
A16-vs-W4A4 relationship.
- W4A4 quality depends on the calibration distribution. This one was calibrated on Russian and
English reranker prompts; a very different domain may behave differently, and the English
regression above suggests the margin is not large.
- Evaluated on two reranking benchmarks. SciDocs documents are paper titles only (~71
characters on average), so it stresses short-document ranking.
- 150 queries per set bounds the resolution of the absolute nDCG figures; the paired Δ columns
are the meaningful comparison.
Compute infrastructure
Quantization — CPU only, no GPU involved. Intel Core i9-14900K (24 cores / 32 threads,
AVX2, no AVX-512 and no AMX), 125 GB RAM, PyTorch 2.13 CPU build, llm-compressor 0.13.0 with
compressed-tensors 0.18.0. Wall time: 104 minutes for 512 calibration samples at 512
tokens, running the sequential pipeline across 37 blocks. The weight-only A16 sibling needs no
calibration at all and takes 24 seconds on the same machine — worth knowing if you plan to
re-quantize rather than download.
Evaluation and serving. 2× NVIDIA RTX 5090 (32 GB, SM120 / compute capability 12.0),
vLLM 0.28.0 on the CUDA 13.0 image, --tensor-parallel-size 2 --enforce-eager. The GPUs were
shared with another resident workload for the entire measurement window, which is why the
throughput figures are presented as a relative ranking rather than as peak performance, why
the KV cache was pinned to 128 MiB, and why the bf16 baseline had to run with
--cpu-offload-gb 2 — 7.48 GiB of weights did not fit in what was free.
License and attribution
Apache-2.0, inherited from
Qwen/Qwen3-Reranker-4B. This repository
contains only quantized weights derived from that model; all model credit belongs to the Qwen
team. Quantized with llm-compressor.