Training Data
Table with columns: Metric, Value| Metric | Value |
|---|
| Pairs | 6,787 |
| Words | 13.4M |
| Selection | Composite gap ≥25, chosen score ≥80 |
| Prompt tokens | 1k-4k (3 length buckets) |
| Completion tokens | ~250 each |
- On-policy: 5 candidates per prompt from writer model via vLLM (T=0.9, top_p=0.92)
- DeepSeek-V4-Pro judge, 5 axes (consistency, garbage, voice, originality, repetition)
- Dominant axes in selected pairs: garbage (54%), repetition (43%)
- Source prompts: fiction continuation from 5,724-sample SFT corpus
Results
Table with columns: Metric, Value| Metric | Value |
|---|
| NLL loss | 2.27 → 2.15 |
| Accuracy | 0.625 → 0.85 |
| Grad norm | 14 avg |
Stopped early — accuracy saturated by step ~125, margins kept growing noisily. Mechanical failure rate: 6.5% pre-ORPO (68,530 gens) → 0% post (160 gens).
logps/rejected cratered (−2 → −13), logps/chosen flat. Model learned what NOT to do.
Training
- 1×H100 80GB SXM (Unsloth + TRL ORPOTrainer)
- LoRA rank 512, alpha 512 (rsLoRA), continued from writer adapter
- 4480 context (measured pair max: 4449 Tekken tokens)
- BF16 base, FP32 AdamW
- ~45 min
Table with columns: Metric, Value| Metric | Value |
|---|
| Steps | 250 / 849 |
| Learning rate | 5e-6 (constant) |
| Batch size | 8 (1 × grad_accum 8) |
| Beta | 0.1 |
Trained with Unsloth + TRL ORPOTrainer.