Evaluated head-to-head against the base model on identical questions, with an 8192-token budget so step-by-step reasoning is never truncated.
Uncensored at no cost to quality. Removing refusals left capability fully intact: it matches the aligned model on MMLU-Pro (88.29%), GSM8K, and hard GPQA-Diamond reasoning (56.57% vs base 57.58%, within noise) - and still beats the base on knowledge.
This model has had refusal behavior reduced, so it may produce content a safety-tuned model would decline. You are responsible for how you use it - use it lawfully and ethically, in accordance with the base-model and dataset licenses. It is not intended for producing harmful, illegal, or abusive content.
Tip: this model reasons before answering - give it a generous max_tokens so it can finish its chain of thought.