Why this exists
NVIDIA ships this model as NVFP4 only. That is the right format for the hardware it targets,
and it is unusable in three situations that are not unusual — each of these is an error, not
an opinion:
1. transformers cannot load it at all.
ValueError: quant_method: "modelopt" ... # no loader in transformers 5.6 or 5.15
So there is no AutoModelForCausalLM path: no LoRA, no finetune, no logit inspection, no
generate() for a quick evaluation. Anything that starts with loading the model in Python
starts by needing these weights instead.
2. NVFP4 needs Blackwell, or a very new vLLM.
ValueError: Current platform does not support NVFP4 quantization. Please use Blackwell andabove. # SGLang, get_min_capability() == 100
vLLM has Hopper W4A16 kernels for it — but only in versions built on torch ≥ 2.11, which is
CUDA 13 on PyPI:
RuntimeError: The NVIDIA driver on your system is too old (found version 12020)
An H100 behind a CUDA 12 driver — which describes a great many university and national-lab
clusters that update drivers on their own schedule — can run this model in bf16 and cannot
run it in NVFP4.
3. Even the right vLLM may not load on an enterprise Linux.
ImportError: /lib64/libm.so.6: version `GLIBC_2.29' not found (required by .../vllm/_moe_C.abi3.so)
RHEL 8 and its rebuilds ship glibc 2.28, and every vLLM wheel builds its MoE kernels against
2.29. vllm/_C needs only 2.14 and loads, so a dense model serves perfectly and a
30B-A3B mixture-of-experts dies partway through its first forward pass. The fix is a
container — which is easier to justify once the weights are in a format the engine inside it
can read without a driver upgrade too.
So: this is for pre-Blackwell hardware, older drivers, enterprise-Linux clusters, and
anybody who wants the model in transformers. If none of those describe you, use
NVIDIA's NVFP4 release:
it is a third of the size and faster on hardware that supports it.
What was changed
Nothing but the number format and the layout.
Table | |
|---|
| routed and shared experts | NVFP4 W4A16 (uint8 nibble pairs, e4m3 group scales, fp32 global scale, group size 16) → bf16 |
mamba in_proj/out_proj | FP8 per-tensor → bf16 |
| everything else | passed through unchanged |
| per-expert 2-D tensors | stacked into the fused [n_experts, out, in] layout transformers expects |
backbone.* names | remapped to model.* |
| MTP layers |
dequant_manifest.json records the source path, shard and byte counts, the tensor count, the
FP4 E2M1 table and the nibble order, so you can confirm you are looking at the same rebuild.
The nibble unpacking is bit-exact against compressed-tensors' reference implementation.
Produced by
dequantize_nemotron_nvfp4.py —
about ten minutes on CPU, ~20 GiB of RAM — so you can also rebuild it yourself from NVIDIA's
checkpoint rather than trusting this upload.
What this is not
- Not a retrain, a finetune or a merge. No gradient step was taken. Benchmarks, intended
use, limitations and the bias, explainability, privacy and safety statements are NVIDIA's
and are unaffected: see
bias.md, explainability.md, privacy.md, safety.md here, and
their model card
for the numbers.
- Not an NVIDIA release. It is a third-party derivative, uploaded by
PursuitOfDataScience, and any mistake in the
conversion is mine rather than theirs.
- Not smaller or faster. It is 66 GB against 21 GB, and on hardware that can run NVFP4 it
is strictly worse. Reach for it only when NVFP4 is not an option.
Licence
Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
The weights remain NVIDIA's under the OpenMDW License Agreement, version 1.1, a copy of
which is in LICENSE and travels with any further distribution, along with the
notices of origin above — which is what that agreement asks of anyone redistributing a
portion of the Model Materials.
Serving it
vllm serve PursuitOfDataScience/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 \ --max-model-len 32768 \ --reasoning-parser nemotron_v3 \ --tool-call-parser qwen3_coder --enable-auto-tool-choice
nemotron_v3 is named on newer vLLM; on 0.11.2 the equivalents are
--reasoning-parser deepseek_r1
(this chat template opens the thinking block itself, so the output carries
</think> and never
<think>) with the same tool parser. An end-to-end deployment of this
checkpoint — engine, OpenAI-compatible gateway and a chat UI — is at
PursuitOfDataScience/manual-api.