PREDICT Arm B SFT
Full-parameter consequence-prediction SFT model for the matched
PREDICT experiment.
- Base:
Qwen/Qwen3-4B-Base
- Data: 374 verified traces from the official MBPP train split
- Mix: 250 direct-success and 124 one-recovery trajectories
- Protocol:
patch -> predict outcome -> KEEP or REVISE -> test -> FINAL
- Training: 5 epochs, 60 optimizer steps, sequence length 768
- Final loss:
0.0334
- Source revision:
8a4089b
An RL-style sampling check (temperature=0.8, top_k=20) passed 74/128
rollouts across 16 training tasks. Fourteen task groups contained both positive
and negative samples, and no rollout hit the length limit. Prediction accuracy
was 84/160 (52.5%). This is a readiness check, not a held-out benchmark.
This checkpoint is the Arm B starting policy for RLVR plus verified-label
prediction CE.