Training
- Method: Group Relative Policy Optimization (GRPO) with LoRA (rank 32), verifiable reward (RLVR), 35 training steps, single-node 4×H200
- Environment: a simulated multi-turn business agent setting with connector-style tool discovery, paginated data access, policy/KB adherence, and programmatically verifiable rewards
- Framework: verl
Results — AutomationBench (60-task subset, self-hosted)
Table with columns: Model, Size, Partial credit, Strict pass| Model | Size | Partial credit | Strict pass |
|---|
| Ornith-1.0-9B | 9B | 22% | 3% |
| seul | 9B | 28% | 7% |
| Qwen3.6-27B | 27B | 34% | 10% |
RLVR training raises partial credit +6 points and more than doubles the strict pass rate over the base, closing a large fraction of the gap to a 3× larger model on the same tasks.