Native MarinSkyRL Tinker-style OPD: one step from SFT step 400
This PEFT LoRA adapter is the result of one full-shape, chosen-token sampled reverse-KL optimizer update in MarinSkyRL. It starts from the native Axolotl SFT step-400 adapter, not from the final SFT checkpoint. Load it on the pinned base model Qwen/Qwen3.5-9B-Base revision 68c46c4b3498877f3ef123c856ecfde50c39f404.
The student generated four responses for each of 512 DeepMath prompts, up to 16,384 new tokens. The chosen-token teacher was Qwen/Qwen3.5-9B revision c202236235762e1c871ad0ccb60c8ee5ba337b9a. The teacher and student used the same tokenizer. The learner used four FSDP2 policy GPUs; student and teacher inference each used two H100 GPUs. The learning rate was 1e-4 and the LoRA rank was 128.
An independent 30-question AIME 2024 evaluation scored 26/30 correct; two answers reached the length cap, so this is not a fully comparable benchmark result. The SFT step-400 starting point scored 19/30 under the same local protocol. The reproduction bundle holds retained evaluation outputs; launch manifests and exact source/config snapshots are in the experiment artifact bundle.
This is a target-score demonstration, not a 200-step trajectory reproduction of the Thinking Machines recipe. The published recipe begins OPD after 3,000 SFT steps. Historical per-token teacher score tensors and student training rollouts were not retained for this gate.
SHA-256: adapter_model.safetensors = c706c786c331bfbab4bb077bb69149b2092e5e7f28544ef531e63d6cb38da071; adapter_config.json = 994763a467f94f8628adf720c4652bd580dd3a2a8bd7c41c70fc80bec70cfd8e.