Architecture
- Base: DeepSeek-R1-Distill-Qwen-1.5B (Qwen2ForCausalLM, 1.5B params)
- LoRA rank 8, alpha 160, on layers 20-27 (q/k/v/o + gate/up/down)
- Merged into a standalone model (no LoRA needed at inference)
Training
Table with columns: Stage, Platform, Steps, Batch, Seq len, Data| Stage | Platform | Steps | Batch | Seq len | Data |
|---|
| MLX run 1 | Apple Silicon | 150 | 4 | 512 | OpenCodeInstruct |
| MLX run 2 | Apple Silicon | 250 | 2 | 1024 | OpenCodeInstruct |
| Phase 1 | 2x RTX 3090 | 550 | 16 (eff) | 2048 | OpenThoughts + OpenR1-Math + OpenCodeInstruct |
Evaluation
Table with columns: Benchmark, Score| Benchmark | Score |
|---|
| GSM8K | 46.0% |
| HumanEval (pass@1) | 7.3% |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('auryn-macmillan/boostedv1')
tok = AutoTokenizer.from_pretrained('auryn-macmillan/boostedv1')
inputs = tok('What is 2+2?', return_tensors='pt')
out = model.generate(**inputs, max_new_tokens=128)
print(tok.decode(out[0]))
Notes
- Standard Qwen2 architecture, no custom code, no trust_remote_code needed.
- Trained in an isolated container; repo contains no training code or data.