Model details
Exact group size, masking ratio and the rest of the configuration are in the launch script linked
below.
Results
ROUGE-L and GPT-4o A/B win rate (%) against the untrained base model.
Table with columns: Benchmark, ROUGE-L (base → SpyRL), A/B win rate| Benchmark | ROUGE-L (base → SpyRL) | A/B win rate |
|---|
| GovReport | 30.2 → 36.7 | 74.6 |
| Multi-News | 23.1 → 26.4 | 80.2 |
| QmSum | 21.3 → 25.3 | 68.4 |
| VcSum | 15.1 → 19.1 | 70.2 |
| SamSum | 43.2 → 48.2 | |
Trained only on GovReport — the gains on the other four benchmarks are zero-shot transfer.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "SpyRL/SpyRL-Qwen3-4B-Summarization"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Your prompt here"}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=1024)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))
The chat template and tokenizer are inherited unchanged from the base model.
Training code
Full training code, launch scripts and per-task environments: https://github.com/wangqinsi1/SpyRL-Self-PlaY-Reinforcement-Learning
git clone https://github.com/wangqinsi1/SpyRL-Self-PlaY-Reinforcement-Learning.git && cd SpyRL-Self-PlaY-Reinforcement-Learning
conda create -n spyrl python=3.10 -y && conda activate spyrl
bash setup.sh
bash spyrl/train_summarization.sh
Citation
@inproceedings{wang2026spyrl,
title = {From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement},
author = {Qinsi Wang and Jing Shi and Huazheng Wang and Kun Wan and Yiran Wu and Bo Liu and Qingyun Wu and Hai Helen Li and Yiran Chen and Handong Zhao and Wentian Zhao},
booktitle = {Conference on Language Modeling (COLM)},
year = {2026}
}