Training
- Base: BoostedV1 Phase-1 (DeepSeek-R1-Distill-Qwen-1.5B + LoRA r=8)
- IL v9: 3 rounds, each generating 600 candidate solutions, externally
verifying, model self-auditing (threshold 5/10), then training 100 steps
on 52-65 high-quality examples per round
- Best round: round_001 (merged here)
Evaluation
Table with columns: Benchmark, BoostedV1 (phase-1), BoostedV1-ILv9| Benchmark | BoostedV1 (phase-1) | BoostedV1-ILv9 |
|---|
| GSM8K | 46.0% | 46.5% |
| HumanEval (pass@1) | 7.3% | 11.0% |
The IL v9 loop produced the largest gains on code generation (HumanEval
pass@1 7.3% -> 11.0%).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('auryn-macmillan/boostedv1-ilv9')
tok = AutoTokenizer.from_pretrained('auryn-macmillan/boostedv1-ilv9')
inputs = tok('Write Python to check if a number is prime.', return_tensors='pt')
out = model.generate(**inputs, max_new_tokens=256)
print(tok.decode(out[0]))
Notes
- Standard Qwen2 architecture, no custom code, no trust_remote_code needed.
- Trained in an isolated container; repo contains no training code or data.