ronpo-llama31-saferlhf-stage3-os-s42
Research checkpoint for the RONPO AAAI revision experiments.
- Method: RONPO OS Stage-3
- Base model:
meta-llama/Llama-3.1-8B-Instruct
- Seed: 42
- Local source at upload time:
/NHNHOME/AIPR/sjkim/MNPO_rev_20260710/results/p7_stage3_fresh_default_test_20260717/stage3/ronpo_os_stage3/train/full
- Uploaded at UTC: 2026-07-17T14:47:27Z
- Run status metadata:
not found
Stage-3 SafeRLHF 1,000-prompt held-out evaluation. The model passed the corrected stability gate. On the fixed panel it ranked first by the preregistered normalized worst-objective point estimate (0.4330; 95% interval [0.4151, 0.4508]); its paired difference versus IPO was 0.0058 with 95% CI [-0.0183, 0.0295], so this is not a statistically significant lead.
Intended use: reproducibility and evaluation for the RONPO paper. This
checkpoint is not intended as a production assistant.