Model Details
- Base Model: Qwen/Qwen3-4B
- Abliteration Method: Heretic v1.2.0
- Trials: 200
- Trial Selected: Trial 96
- Refusals: 3/100 (vs 100/100 original)
- KL Divergence: 0.0000 (zero measurable model damage)
Files
model-00001-of-00002.safetensors
model-00002-of-00002.safetensors
config.json
tokenizer.json
tokenizer_config.json
ComfyUI Format (for Z-Image / FLUX.2 Klein 4B text encoder)
comfyui/qwen3-4b-heretic.safetensors # bf16, 7.5GB
comfyui/qwen3-4b-heretic_fp8_e4m3fn.safetensors # fp8 row-wise, 4.2GB
comfyui/qwen3-4b-heretic_int8.safetensors # int8 ConvRot row-wise, 4.2GB
comfyui/qwen3-4b-heretic_int4.safetensors # int4 W4A4 ConvRot, 2.5GB
comfyui/qwen3-4b-heretic_nvfp4.safetensors # nvfp4, 2.7GB
comfyui/qwen3-4b-heretic_mxfp8.safetensors # mxfp8, 4.3GB
Quality: All quantized variants use SVD-guided learned rounding (AdaRound via convert_to_quant), which optimizes each weight's rounding direction to minimize output reconstruction error — noticeably higher fidelity than naive round-to-nearest quantization.
Table with columns: Quant, Size, Notes| Quant | Size | Notes |
|---|
| F16 | ~7.5GB | Lossless reference |
| Q8_0 | ~4GB | Excellent quality |
| Q6_K | ~3GB | Very good quality |
| Q5_K_M | ~2.7GB | Good quality |
| Q4_K_M | ~2.3GB | Recommended balance |
| Q3_K_M | ~1.9GB |
All variants load natively in ComfyUI 0.30.0+ (no plugins) via the comfy_quant metadata embedded in each file.
Table with columns: Format, Size, Bits, Notes| Format | Size | Bits | Notes |
|---|
| FP8 (E4M3, row-wise) | 4.2GB | 8 | Best speed/quality balance; works on Ada/Hopper+ |
| INT8 (ConvRot row-wise) | 4.2GB | 8 | Hadamard-rotated; broad GPU support |
| MXFP8 | 4.3GB | 8 | Microscaling FP8 (E8M0 block scales); Blackwell-accelerated |
| INT4 (W4A4 ConvRot) | 2.5GB | 4 |
NVFP4/MXFP8 inference is fastest on Blackwell (RTX 5090/5080, SM100+), but ComfyUI also supports software dequantization on older GPUs (tested working on RTX 4090). INT8 and INT4 both use Hadamard rotation (ConvRot); INT4 W4A4 uses ComfyUI's convrot_w4a4 path.
Usage
With ComfyUI (Z-Image / FLUX.2 Klein 4B)
-
Download a ComfyUI format file:
- FP8 (recommended):
comfyui/qwen3-4b-heretic_fp8_e4m3fn.safetensors (4.2GB)
- INT4 (smallest):
comfyui/qwen3-4b-heretic_int4.safetensors (2.5GB)
- NVFP4:
comfyui/qwen3-4b-heretic_nvfp4.safetensors (2.7GB)
- INT8:
comfyui/qwen3-4b-heretic_int8.safetensors (4.2GB)
- MXFP8:
comfyui/qwen3-4b-heretic_mxfp8.safetensors (4.3GB)
- bf16 (full precision):
comfyui/qwen3-4b-heretic.safetensors (7.5GB)
-
Place in
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"DreamFast/qwen3-4b-heretic",
device_map="auto",
torch_dtype=torch.bfloat16
)
tokenizer = AutoTokenizer.from_pretrained("DreamFast/qwen3-4b-heretic")
prompt = "Describe a dramatic sunset over a cyberpunk city"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
With llama.cpp
llama-server -m qwen3-4b-heretic-Q4_K_M.gguf
Abliteration Process
Created using Heretic v1.2.0 with 200 optimization trials:
? Which trial do you want to use?
> [Trial 96] Refusals: 3/100, KL divergence: 0.0000 <-- selected
[Trial 90] Refusals: 5/100, KL divergence: 0.0000
[Trial 95] Refusals: 9/100, KL divergence: 0.0000
[Trial 122] Refusals: 90/100, KL divergence: 0.0000
...
Trial 96 was selected for having the fewest refusals (3/100) with zero measurable KL divergence, indicating the abliteration surgically removed the refusal mechanism with no damage to model capabilities.
Limitations
- This model inherits all limitations of the base Qwen 3 4B model
- Abliteration reduces but does not completely eliminate refusals (3/100 remain)
License
This model is released under the Apache 2.0 License, following the base Qwen 3 4B model license.
Acknowledgments