Model Details
- Base Model: Qwen/Qwen3-4B-Thinking-2507 (a reasoning model)
- Training Method: GRPO (Group Relative Policy Optimization), full fine-tuning
- Reward Model: idealab-cs2/reappraisal-bt-reward-model
- Training Data: scenarios from Li et al. (2025) — 6 negative interpersonal vignettes
- Hardware: 4× H100, DeepSpeed ZeRO-2
- Trained By: ruggsea
Model Description
The model was trained with the following system prompt:
You are helping someone reduce a negative emotion they feel in a short interpersonal scenario byoffering an alternative interpretation of the situation (a 'reappraisal' / 'rethinking'). Directyour response at the person in the scenario in second person. Limit your response to two sentencesmaximum. Do not list emotions; output only the reappraisal text.
The user turn is
SCENARIO:\n{scenario}\n\nWrite a reappraisal of this scenario in two sentences maximum, addressing the person in second person.
Because the base is a thinking model, generations
contain a reasoning trace in
<think>...</think> followed by the reappraisal; only the text after
</think> is the answer (and the only part scored by the reward during training).
The reward is the Bradley-Terry reward model's scalar score of the answer, with a format gate that
penalizes outputs that are empty, longer than two sentences, or over ~340 characters (to prevent
the policy from inflating the reward through verbosity). The KL penalty to the base model is the
main regularizer.
Evaluation
Reappraisals were compared pairwise against GPT-4-0314's reappraisals for the same scenarios, judged
by Llama-3.1-70B-Instruct. That judge was selected by calibration against the human preference data
(it agrees with the human-preferred reappraisal 93% of the time on this task; larger judges such as
Llama-3.1-405B-FP8 and GLM-5.2 agreed less and showed position bias).
Table with columns: model, win-rate vs GPT-4-0314, n, 95% CI lower| model | win-rate vs GPT-4-0314 | n | 95% CI lower |
|---|
| Reappraisal-4B-GRPO (this model) | 0.581 | 480 | 0.537 |
| Qwen3-4B-Thinking (base, untrained) | ~0.46 | — | — |
| Qwen3-4B-Instruct + same GRPO recipe (non-reasoning) | 0.458 | 120 | 0.369 |
Training hyperparameters
- learning_rate: 8e-6
- beta (KL coefficient): 0.04
- num_generations: 8
- max_steps: 250
- per_device_train_batch_size: 8, gradient_accumulation_steps: 4
- max_completion_length: 1280
- precision: bf16, DeepSpeed ZeRO-2 across 4 GPUs
Usage
import torchfrom transformers import AutoModelForCausalLM, AutoTokenizer tok = AutoTokenizer.from_pretrained("idealab-cs2/reappraisal-4b-grpo")model = AutoModelForCausalLM.from_pretrained( "idealab-cs2/reappraisal-4b-grpo", torch_dtype=torch.bfloat16, device_map="cuda") system = ("You are helping someone reduce a negative emotion they feel in a short interpersonal " "scenario by offering an alternative interpretation of the situation (a 'reappraisal' / " "'rethinking'). Direct your response at the person in the scenario in second person. " "Limit your response to two sentences maximum. Do not list emotions; output only the " "reappraisal text.")scenario = "A coworker took credit for your idea in a meeting and you feel humiliated."user = f"SCENARIO:\n{scenario}\n\nWrite a reappraisal of this scenario in two sentences maximum, addressing the person in second person." ids = tok.apply_chat_template([{"role": "system", "content": system}, {"role": "user", "content": user}], add_generation_prompt=True, return_tensors="pt").to(model.device)out = model.generate(ids, max_new_tokens=1280, do_sample=True, temperature=0.7, top_p=0.9)text = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)reappraisal = text.rsplit("</think>", 1)[-1].strip() # answer follows the reasoning traceprint(reappraisal)
References