Results (dev_public, 300 puzzles, greedy, exact verifier)
Table with columns: model, accuracy, hard (5-num), format_rate, avg_tokens (correct)| model | accuracy | hard (5-num) | format_rate | avg_tokens (correct) |
|---|
| base Qwen2.5-0.5B-Instruct (floor) | 0.33% | 0.00% | 0.00% | 20.0 |
| this model (GRPO, lr 3e-6) | 12.00% | 1.67% | 0.00% | 16.9 |
36x over the base floor (1/300 -> 36/300). Ablations: learning rate was the only lever
that moved accuracy (1e-6 -> 3e-6 gave 1% -> 12%); more steps (3000) and larger group (16)
both plateaued. Known failure mode: reasoning collapse — the shaped reward scores only the
<answer>, so the model emits bare answers with no <think> (format_rate = 0).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")tok = AutoTokenizer.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")
Trained for the RLVR Arena capstone (RL in Production Bootcamp).