What this is
The model has had its refusal direction ablated via a difference-of-means LoRA on the residual stream, then merged back into the base weights. The goal is to remove knee-jerk refusals while preserving the base model's reasoning capability.
Original model
This is a derivative of WeiboAI/VibeThinker-3B. Go take a look at the original repo for full details on the base — architecture, training procedure, intended use cases, and licensing terms all carry over from there.
Abliteration metrics
Table with columns: Metric, This model, Original (by definition)| Metric | This model | Original (by definition) |
|---|
| KL divergence | 0.0870 | 0 |
| Refusals | 2/100 | ~64/100 |
KL divergence of 0.087 is low — well under the 0.5 threshold above which abliteration typically starts damaging model capability. The 2/100 refusal count on the held-out eval set indicates the abliteration generalized rather than just memorizing training prompts.
How it was made
- Tool: Heretic v1.4.0 (
heretic-llm)
- Method: Refusal direction computed as
mean(bad_residuals) - mean(good_residuals) per layer, projected out of attention and MLP weights via a rank-3 LoRA adapter, then merged
- Training prompts: Custom refusal-triggering set (one prompt per line, plain text, special themes)
- Good/harmless set:
mlabonne/harmless_alpaca (train[:400])
- Eval set:
mlabonne/harmful_behaviors (test[:100])
- System prompt:
You are a helpful assistant.
To reproduce the same pipeline on any other model:
pip install -U heretic-llmheretic --model WeiboAI/VibeThinker-3B
Then follow the interactive prompts. See the Heretic repo for the full parameter space.
Intended use
Same as the base model, minus refusal behavior. Apply your own application-level review or safety layer for any deployment where end-user-facing safety matters — this model ships with none.
It has been specifically designed to be fully uncensored across all topics. Please use it responsibly and with care.
Bias, risks, and limitations
- All limitations of the base model apply.
- The model will comply with requests the base model would have refused. Use accordingly.
- Abliteration is approximate. A small fraction of refusals may persist (here: ~2/100), and some unrelated capabilities may shift slightly. The KL divergence figure is your best signal for how much the abliteration perturbed the base distribution.
- Not affiliated with, endorsed by, or derived from any Anthropic product despite any naming similarities in third-party derivatives.
License
Inherits from WeiboAI/VibeThinker-3B. Check the original repo for the exact terms.
Credits