Model Details
- Base Model: Qwen/Qwen3-4B-Thinking-2507 (a reasoning model; non-reasoning bases lose here)
- Training Method: GRPO (Group Relative Policy Optimization), full fine-tuning, DeepSpeed ZeRO-2, 4× H100
- Reward Model: idealab-cs2/reappraisal-reward-model-v2 (3-seed ensemble, reward =
mean − 0.5·std)
- Training Data: scenarios from Li et al. (2025) — 6 negative interpersonal vignettes
- Trained By: ruggsea
Evaluation
Reappraisals were compared pairwise against GPT-4-0314's reappraisals for the same scenarios, judged
by Llama-3.1-70B-Instruct scoring both A/B orderings. That judge was selected by calibration against
the human preference data (it agrees with the human-preferred reappraisal 93% of the time on this
task; larger judges such as Llama-3.1-405B-FP8 and GLM-5.2 agreed less and showed position bias).
Results are pooled over three independent evals of this checkpoint (fresh generations + an
independent both-orders judge each), and marginal wins are confirmed at n ≥ 240 with the lower CI
bound above the bar.
Table with columns: model, win-rate vs GPT-4-0314, n, 95% CI lower (iid, optimistic)| model | win-rate vs GPT-4-0314 | n | 95% CI lower (iid, optimistic) |
|---|
| Reappraisal-4B-GRPO-RMv2 (this model) | 0.806 | 720 | 0.776 |
| Reappraisal-4B-GRPO (v1, BT reward) | 0.581 | 480 | 0.537 |
| Qwen3-4B-Thinking (base, untrained) | ~0.46 | — | — |
| DeepSeek-R1-671B (single-pass reference) | 0.771 | 240 | 0.718 |
Read the CI column as a screening tool, not a precise bound. The 720 contests cluster on only
6 vignettes (n_effective ≈ 6), so the iid Wilson CIs above are optimistic, and the earlier claim
that "the CI-low 0.776 sits above the R1 reference 0.771" is retracted — comparing an iid CI-low to
a reference point is not a valid test. The recipe is seed-robust (a different-seed run scored 0.854),
and the win over GPT-4-0314 on these 6 vignettes is real and clears the 0.608 bar with margin; but this
is a specialist trained on 6 scenarios, and the strongest policy in this project is the
committee model (0.872 6-vignette,
vignette-cluster CI-low 0.781), which is the checkpoint to read for the honest R1 comparison.
Out-of-distribution: this does NOT generalize past parity, and no 4B in this project beats R1
off-distribution. The policy is GRPO-trained on the same 6 vignettes it is evaluated on. The earlier
"0.784 off-vignette, well above 0.5" claim is retracted: it used a weaker gemma opponent and iid
statistics. Against a strong open 72B on fresh cross-family scenarios, our best policies are at
parity-or-below (cluster CIs spanning 0.5), and against DeepSeek-R1-671B off-distribution our 4B
loses — single-pass ~0.39–0.45 and best-of-N only tied on the selection pool before reverting to
~0.375 on a fresh held-out pool (winner's curse). Off-distribution generalization tracks base capacity,
not our training.
The improvement over the v1 policy comes almost entirely from the reward model: the same GRPO recipe
against the earlier reward models scored 0.581 (Bradley-Terry) and 0.698 (multidimensional), against
0.806 with the v2 ensemble.
Prompting
The model was trained with this system prompt:
You are helping someone reduce a negative emotion they feel in a short interpersonal scenario byoffering an alternative interpretation of the situation (a 'reappraisal' / 'rethinking'). Directyour response at the person in the scenario in second person. Limit your response to two sentencesmaximum. Do not list emotions; output only the reappraisal text.
The user turn is
SCENARIO:\n{scenario}\n\nWrite a reappraisal of this scenario in two sentences maximum, addressing the person in second person.
Because the base is a thinking model, generations
contain a reasoning trace in
<think>...</think> followed by the reappraisal; only the text after
</think> is the answer.
Training hyperparameters
- learning_rate: 8e-6
- beta (KL coefficient): 0.04 (kept strictly > 0; β=0 collapses and reward-hacks)
- num_generations: 12
- max_steps: 500
- precision: bf16, DeepSpeed ZeRO-2 across 4 GPUs
The reward is the ensemble score of the answer with a format gate penalizing outputs that are empty
or longer than two sentences (so the policy cannot inflate reward through verbosity). Training was
healthy: reward rising, KL bounded, entropy alive, no reward-std collapse (wandb vzo3wpv6).
Data separation
The 6 vignettes' GPT-4-0314 reappraisals and the frontier-opponent answers are held out and used only
for evaluation; training uses the vignette prompts (the task) and the human-rating-based reward model.
Any pooled or synthetic training text is n-gram-checked against the held-out eval answers before use.
Intended use & limitations
Research artifact for the study of RLHF, reward modeling, and computational emotion regulation. It
writes short cognitive reappraisals of negative interpersonal situations. It is not a clinical or
mental-health tool and must not be used as one. Quality is validated only on the narrow reappraisal
task and the evaluation above.
References