Why this checkpoint is published
It is the paired no-exploration control for
ngqtrung/video-8b-grpo-n16-warm-explore,
the campaign's highest score. The pair is the whole point: same corpus, same group size, same step,
one variable.
Table with columns: core-3 @40 | core-3 @40 |
|---|
| n16 + exploration (published sibling) | 0.4942 |
| this control, no exploration | 0.4879 |
| difference | +0.63 pt for exploration |
That 0.63 pt sits inside the campaign's own noise — the top five checkpoints span 1.0 pt on the same
5,645 rows. Publishing the control is what lets anyone check that for themselves rather than taking
the winner's margin at face value.
Evaluation
core-3 = 5,645 rows, num_failed = 0.
Table with columns: mean_accuracy, VideoMME, Video-Holmes, PerceptionComp | mean_accuracy | VideoMME | Video-Holmes | PerceptionComp |
|---|
| gs40 | 0.4879 | 0.6570 | 0.4692 | 0.3375 |
| stock base | 0.4426 | 0.6404 | 0.4066 | 0.2807 |
Protocol caveat. Measured before 2026-08-23, without <think> prefill at eval time, so format
reads 0.0. Internally consistent; do not mix with prefill-era numbers.
Training setup
Table | |
|---|
| algorithm | GRPO, fully-async (verl fork), no exploration |
| corpus | video v3 24f100k, multiple-choice video QA |
| group size | rollout.n = 16 |
| topology | 4 nodes, 2 trainer + 2 rollout |