Introduction
Pyrex 8B Instruct Uncensored is a real, trained model — built with our own
data recipe, our own QLoRA fine-tune, and our own evaluation pipeline on
Hugging Face GPU compute. It is a coding-first instruction model designed for
real developer work: writing code, fixing bugs, explaining technical problems,
and driving agent loops.
Key strengths
- Coding-first — trained mostly on high-quality code instructions
(OpenCoder real-user data, CodeAlpaca, evol-codealpaca), so it writes
clean and correct code.
- Uncensored — answers directly and honestly, without refusals.
- Agentic / tool-use ready — function-calling data was part of training;
suited for OpenAI-compatible APIs and agent workflows.
- Portable — full-precision safetensors plus GGUF quants for Ollama,
llama.cpp and LM Studio.
Uncensored usage note. This model has no safety-alignment softening built
in. It can produce content that is explicit, offensive, or otherwise
unsuitable for some audiences. Use it responsibly and at your own discretion —
the model and its authors assume no liability for outputs.
Model details
Table with columns: Property, Value| Property | Value |
|---|
| Parameters | 7.6B |
| Architecture | Transformers · RoPE · SwiGLU · RMSNorm · GQA |
| Context window | Up to 32k native (trained at 3072) |
| Base | Coding-specialized instruct model (uncensored variant) |
| Method | QLoRA (4-bit NF4) + LoRA r=48, α=96 |
| Chat format | chatml (`< |
| License | Apache-2.0 |
Benchmarks
Measured with greedy pass@1 on HumanEval — 164 unseen problems, official
unit tests executed in a sandbox. No sampling luck, no leaked answers.
Table with columns: Model, HumanEval pass@1| Model | HumanEval pass@1 |
|---|
| Base (untuned) | 29.3% (48/164) |
| Pyrex (previous) | 41.5% (68/164) |
| Pyrex 8B Instruct Uncensored | 52.4% (86/164) |
+23.1 points over the base on unseen problems.
Usage
Requirements
transformers >= 4.44 for best apply_chat_template support.
- Any recent PyTorch build (CUDA or CPU).
Python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "ImposterOnline/Pyrex-8B-Instruct-Uncensored"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto", torch_dtype="auto")
messages = [{"role": "user", "content": "Explain async/await in Python."}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
out = model.generate(**tok(text, return_tensors="pt"), max_new_tokens=512)
print(tok.decode(out[0], skip_special_tokens=True))
Chat format is chatml. System prompt is optional; when used it is a system
message before the user turn.
Server / agentic
Works with any OpenAI-compatible server (vLLM, llama.cpp server, Ollama,
Text Generation Inference). Function calling was part of training.
vllm serve ImposterOnline/Pyrex-8B-Instruct-Uncensored
GGUF
Quantized builds live in
Pyrex-8B-Instruct-Uncensored-GGUF.
Table with columns: File, Size, Notes| File | Size | Notes |
|---|
pyrex-q4_k_m.gguf | 4.4 GB | Fast, recommended daily driver |
pyrex-q5_k_m.gguf | 5.1 GB | Higher quality |
pyrex-q8_0.gguf | 8.1 GB | Fast, best quality |
pyrex-f16.gguf | 15.2 GB | Full precision |
# Ollama (Modelfile ships in the GGUF repo)
ollama create pyrex-8b -f Modelfile
ollama run pyrex-8b "Write a Python function that merges overlapping intervals."
# llama.cpp
./llama-cli -m pyrex-q4_k_m.gguf \
-p "<|im_start|>user\nYour question here\n<|im_end|>\n<|im_start|>assistant\n" \
-n 512 -c 8192
Training
Table with columns: Parameter, Value| Parameter | Value |
|---|
| Data | ~95k cleaned examples (code + tool-use + general) |
| Sequence length | 3072 |
| Loss masking | Assistant-only (completion-only SFT) |
| Optimizer | AdamW · lr 2e-4 · cosine · warmup 3% |
| Epochs | 1 |
| Hardware | Hugging Face A100 GPU job |
Data sources (credited): OpenCoder-LLM/opencoder-sft-stage1
(realuser + largescale-diverse), sahil2801/CodeAlpaca-20k,
theblackcat102/evol-codealpaca-v1, databricks/databricks-dolly-15k,
NousResearch/hermes-function-calling-v1.
The pipeline is config-driven and reproducible — prepare_data.py →
train_qlora.py → eval_humaneval.py → publish.sh — in the companion
repo Arhan-w/pyrut.
License
Apache-2.0.