Model Overview
- Model Architecture: Gemma 4
- Input: Text / Image
- Output: Text
- Model Optimizations:
- Weight quantization: FP4
- Activation quantization: FP4
- Release Date: 2026-07-02
- Version: 1.0
- Quantized by: xdavxd
- Base Model: coder3101/gemma-4-31B-it-heretic
- Original Model: google/gemma-4-31B-it
This model is a quantized version of coder3101/gemma-4-31B-it-heretic.
It was evaluated on several tasks to assess its quality in comparison to the original model.
Model Optimizations
This model was obtained by quantizing the weights and activations of coder3101/gemma-4-31B-it-heretic to FP4 data type using the NVFP4 format, ready for inference with vLLM.
This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 67%.
Weights are quantized with FP4 (group_size=16), and activations are quantized with FP4 using local per-group scaling. Only the weights and activations of the linear operators within transformer blocks are quantized using LLM Compressor. Vision tower, embedding, and output head layers are kept in their original precision.
Deployment
Use with vLLM
This model can be deployed using vLLM.
For detailed instructions including multi-GPU deployment, multimodal inference, thinking mode, function calling, and benchmarking, see the Gemma 4 vLLM usage guide.
- Start the vLLM server:
vllm serve xdavxd/gemma-4-31B-it-heretic-NVFP4 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90
To enable thinking/reasoning and tool calling:
vllm serve xdavxd/gemma-4-31B-it-heretic-NVFP4 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90 \
--enable-auto-tool-choice \
--reasoning-parser gemma4 \
--tool-call-parser gemma4 \
--chat-template examples/tool_chat_template_gemma4.jinja \
--limit-mm-per-prompt '{"image": 4, "audio": 1}' \
--async-scheduling
Tip: For text-only workloads, pass --limit-mm-per-prompt '{"image": 0, "audio": 0}' to skip vision encoder memory allocation and free up GPU memory for a longer context window.
- Send requests to the server:
from openai import OpenAI
openai_api_key = "EMPTY"
openai_api_base = "http://<your-server-host>:8001/v1"
client = OpenAI(
api_key=openai_api_key,
base_url=openai_api_base,
)
model = "xdavxd/gemma-4-31B-it-heretic-NVFP4"
messages = [
{"role": "user", "content": "Explain quantum mechanics clearly and concisely."},
]
outputs = client.chat.completions.create(
model=model,
messages=messages,
)
generated_text = outputs.choices[0].message.content
print(generated_text)
Creation
This model was created by applying NVFP4 quantization with LLM Compressor, as presented in the code snippet below.
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
MODEL_ID = "./gemma-4-31B-it-heretic"
SAVE_DIR = "gemma-4-31B-it-heretic-NVFP4"
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
recipe = QuantizationModifier(
targets="Linear",
scheme="NVFP4",
ignore=["re:.*vision.*", "re:.*audio.*", "lm_head", "re:.*embed.*"],
)
oneshot(
model=model,
tokenizer=tokenizer,
recipe=recipe,
dataset="ultrachat_200k",
splits="train_sft",
num_calibration_samples=512,
max_seq_length=2048,
)
model.save_pretrained(SAVE_DIR, save_compressed=True)
tokenizer.save_pretrained(SAVE_DIR)
Evaluation
This model was evaluated on GSM8K Platinum, MMLU-Pro, IFEval, MATH-500, AIME 2025, GPQA Diamond, and LiveCodeBench v6 using lm-evaluation-harness and lighteval, served with vLLM (OpenAI-compatible API). All evaluations were performed with thinking disabled (chat template default).
(as reported by coder3101)
Table with columns: Metric, This model, Original model (google/gemma-4-31B-it)| Metric | This model | Original model (google/gemma-4-31B-it) |
|---|
| KL divergence | 0.0434 | 0 (by definition) |
| Refusals | 15/100 | 99/100 |
Accuracy
Original and RedHatAI benchmarks taken from RedHatAI, comparison to see how the heretic model converted to NVFP4 stacks up against the non-heretic versions. Recovery vs. Original to see how the combination of heretic ARA (Arbitrary-Rank Ablation) and then quantization affects accuracy.
Reproduction
The results were obtained using the following commands:
Each benchmark was run 3 times with different judgmental seeds (1234, 2345, 3456) and the scores were averaged; AIME 2025 used 8 seeds.
Ran using a docker image, you will need to tweak these for your own environment. Kept parameters similar to the RedHatAI runs.
vLLM server (all benchmarks):
sudo docker run --rm -it \
--gpus all \
--ipc=host \
--network host \
-v ~/quant-workspace/gemma-4-31B-it-heretic-NVFP4:/models/heretic:ro \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-e VLLM_USE_V2_MODEL_RUNNER=1 \
ghcr.io/timothystewart6/vllm-gb10:latest \
vllm serve /models/heretic \
--host 0.0.0.0 --port 8001 \
--served-model-name heretic \
--max-model-len 69632 \
--gpu-memory-utilization 0.50 \
--enable-auto-tool-choice \
--reasoning-parser gemma4 \
--tool-call-parser gemma4 \
--limit-mm-per-prompt '{"image":0,"audio":0}' \
--kv-cache-dtype fp8 \
--max-num-seqs 32 \
--async-scheduling
GSM8K Platinum (lm-eval, 0-shot, 3 repetitions)
for SEED in 1234 2345 3456; do
lm_eval --model local-chat-completions \
--tasks gsm8k_platinum_cot_llama \
--model_args "model=heretic,max_length=36096,base_url=http://0.0.0.0:8001/v1/chat/completions,num_concurrent=32,max_retries=3,tokenized_requests=False,tokenizer_backend=None,timeout=2400" \
--num_fewshot 0 \
--apply_chat_template \
--output_path "results_gsm8k_seed${SEED}.json" \
--seed $SEED \
--gen_kwargs "do_sample=True,temperature=1.0,top_p=0.95,top_k=64,max_gen_toks=32000,seed=${SEED}"
echo "Seed $SEED complete"
done
MMLU-Pro (lm-eval, 0-shot, 3 repetitions)
for SEED in 1234 2345 3456; do
lm-eval run --model local-chat-completions \
--tasks mmlu_pro \
--model_args "model=heretic,max_length=36096,base_url=http://0.0.0.0:8001/v1/chat/completions,num_concurrent=32,max_retries=3,tokenized_requests=False,tokenizer_backend=None,timeout=2400" \
--num_fewshot 0 \
--apply_chat_template \
--output_path "results_mmlu_pro_seed${SEED}.json" \
--seed $SEED \
--gen_kwargs "do_sample=True,temperature=1.0,top_p=0.95,top_k=64,max_gen_toks=32000,seed=${SEED}"
echo "MMLU-Pro seed $SEED complete"
done
IFEval (lm-eval, 0-shot, 3 repetitions)
for SEED in 1234 2345 3456; do
lm_eval --model local-chat-completions \
--tasks ifeval \
--model_args "model=heretic,max_length=36096,base_url=http://0.0.0.0:8001/v1/chat/completions,num_concurrent=32,max_retries=3,tokenized_requests=False,tokenizer_backend=None,timeout=2400" \
--num_fewshot 0 \
--apply_chat_template \
--output_path "results_ifeval_seed${SEED}.json" \
--seed $SEED \
--gen_kwargs "do_sample=True,temperature=1.0,top_p=0.95,top_k=64,max_gen_toks=32000,seed=${SEED}"
echo "IFEval seed $SEED complete"
done
MATH-500, AIME 2025, GPQA Diamond (lighteval, 3 repetitions; 8 for AIME 2025)
litellm_config.yaml:
model_parameters:
provider: hosted_vllm
model_name: hosted_vllm/heretic
base_url: http://0.0.0.0:8001/v1
api_key: ''
timeout: 7200
concurrent_requests: 32
generation_parameters:
temperature: 1.0
max_new_tokens: 65536
top_p: 0.95
top_k: 64
seed: 1234
# GPQA Diamond (3 seeds)
for SEED in 1234 2345 3456; do
sed -i "s/seed: .*/seed: $SEED/" litellm_config.yaml
lighteval endpoint litellm litellm_config.yaml 'gpqa:diamond|0' \
--output-dir "results_gpqa_seed${SEED}/" --save-details
echo "GPQA Diamond seed $SEED complete"
done
# MATH-500 (3 seeds)
for SEED in 1234 2345 3456; do
sed -i "s/seed: .*/seed: $SEED/" litellm_config.yaml
lighteval endpoint litellm litellm_config.yaml 'math_500|0' \
--output-dir "results_math500_seed${SEED}/" --save-details
echo "MATH-500 seed $SEED complete"
done
# AIME 2025 (8 seeds)
for SEED in 1234 2345 3456 4567 5678 6789 7890 8901; do
sed -i "s/seed: .*/seed: $SEED/" litellm_config.yaml
lighteval endpoint litellm litellm_config.yaml 'aime25|0' \
--output-dir "results_aime_seed${SEED}/" --save-details
echo "AIME 2025 seed $SEED complete"
done
LiveCodeBench v6 (lighteval, 3 repetitions)
litellm_config.yaml:
model_parameters:
provider: hosted_vllm
model_name: hosted_vllm/heretic
base_url: http://0.0.0.0:8001/v1
api_key: ''
timeout: 7200
concurrent_requests: 32
generation_parameters:
temperature: 1.0
max_new_tokens: 32768
top_p: 0.95
top_k: 64
seed: 1234
for SEED in 1234 2345 3456; do
sed -i "s/seed: .*/seed: $SEED/" litellm_config.yaml
lighteval endpoint litellm litellm_config.yaml 'lcb:codegeneration_v6|0' \
--output-dir "results_lcb_seed${SEED}/" --save-details
echo "LiveCodeBench v6 seed $SEED complete"
done