Metric: val-aux/sciknoweval/reward/mean@16 from 10-step validation logs.
Table with columns: best mean@16, best step, final mean@16, final step| best mean@16 | best step | final mean@16 | final step |
|---|
| 79.49% | 100 | 79.49% | 100 |

Table with columns: step, mean@16| step | mean@16 |
|---|
| 10 | 45.54% |
| 20 | 59.64% |
| 30 | 67.20% |
| 40 | 69.05% |
| 50 | 72.11% |
| 60 | 73.90% |
| 70 | 74.58% |
| 80 | 76.52% |
| 90 | 76.43% |
| 100 | 79.49% |
Files included with this repo:
metrics.json: parsed validation summary
eval_mean16.csv: step-level validation curve data
eval_mean16.png: validation curve plot
Important: the uploaded weights are the final global_step_100/actor checkpoint. If best step is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally.
Checkpoint source:
/mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/chemistry/qwen3gen-chemistry-GRPO-Qwen-Qwen3-8B-mbs8-train64-rollout8-lr1e-6-vllm0.8
W&B run: run-20260630_103227-flhubvrx
This upload uses global_step_100/actor converted from VERL FSDP shards to Hugging Face format.