Trained with LLaMA-Factory full SFT, DeepSpeed ZeRO-3, BF16, and FlashAttention-2.
Table with columns: Hyperparameter, Value
Hyperparameter
Value
Learning rate
7e-6
LR schedule
cosine, warmup ratio 0.1
Epochs
3
Global batch size
128
Max sequence length
4096
Weight decay
0.1
Precision
BF16
Seed
42
Final train loss: 0.6852.
Intended use
Research checkpoint for ARPO-style agentic SFT. Not evaluated as a standalone chat model. Follow the licenses of the base model and the SFT dataset.
Citation
bibtex
@article{dong2025arpo,
author = {Guanting Dong and Hangyu Mao and Kai Ma and Licheng Bao and Yifei Chen and Zhongyuan Wang and Zhongxia Chen and Jiazhen Du and Huiyang Wang and Fuzheng Zhang and Guorui Zhou and Yutao Zhu and Ji-Rong Wen and Zhicheng Dou},