Why this checkpoint is published
It is the losing arm of a tau ablation, published so the rule it produced can be checked.
Lowering the exploration confidence threshold from 0.95 to 0.85 costs accuracy, and this is the
video-domain evidence for the campaign rule "keep tau >= 0.9".
Its tau-0.95 sibling is
ngqtrung/video-8b-grpo-twowave-gs140
at 0.4853.
Evaluation — read the row count before using this number
Table with columns: mean_accuracy, rows, VideoMME, Video-Holmes, PerceptionComp | mean_accuracy | rows | VideoMME | Video-Holmes | PerceptionComp |
|---|
| gs40 | 0.4732 | 4,838 | 0.6625 (n=2,314) | 0.4625 (n=1,574) | 0.2947 (n=950) |
| stock base | 0.4426 | 5,645 | 0.6404 | 0.4066 | 0.2807 |
This score was measured on a 4,838-row subset, not the full 5,645-row core-3 set.
num_failed = 0, so the file looks entirely normal and gives no warning — the row count is the only
signal. It is therefore not directly comparable to the 5,645-row numbers on the sibling cards.
Treat 0.4732 as indicative, not as a ranked result.
Protocol caveat. Also measured before 2026-08-23, without <think> prefill at eval time, so
format reads 0.0.
Training setup
Table | |
|---|
| algorithm | GRPO, fully-async (verl fork) + two-wave token-dropout exploration |
| exploration | two-wave both-degenerate block, tau 0.85 |
| corpus | video v3 24f100k, multiple-choice video QA |
| topology | 4 nodes, 2 trainer + 2 rollout |