Why this checkpoint is published
It is the tau-0.95 arm of a tau sweep, and it is the evidence for the campaign's standing rule
"keep tau >= 0.9". Its tau-0.85 sibling is
ngqtrung/video-8b-grpo-twowave-tau085-gs40.
Table with columns: arm, core-3 peak| arm | core-3 peak |
|---|
| two-wave tau 0.95 (this) | 0.4853 @140 |
| two-wave tau 0.85 (sibling) | 0.4732 @40 (n=4,838 subset — see that card) |
| stock base | 0.4426 |
Until this run, "keep tau >= 0.9" rested only on image-domain (OMR) data. This is the video-side
corroboration.
Evaluation
core-3 = 5,645 rows, num_failed = 0.
Table with columns: mean_accuracy, VideoMME, Video-Holmes, PerceptionComp | mean_accuracy | VideoMME | Video-Holmes | PerceptionComp |
|---|
| gs140 | 0.4853 | 0.6504 | 0.4725 | 0.3330 |
| stock base | 0.4426 | 0.6404 | 0.4066 | 0.2807 |
Protocol caveat. Measured before 2026-08-23, without <think> prefill at eval time, so format
reads 0.0. Internally consistent; do not mix with prefill-era numbers.
Training setup
Table | |
|---|
| algorithm | GRPO, fully-async (verl fork) + two-wave token-dropout exploration |
| exploration | two-wave both-degenerate block, tau 0.95 |
| corpus | video v3 24f100k, multiple-choice video QA |
| topology | 4 nodes, 2 trainer + 2 rollout |