Recipe
- Base model:
Qwen/Qwen3-VL-4B-Instruct
- Adapter: LoRA, rank 32, alpha 32 (
task_type: CAUSAL_LM)
- Algorithm: GRPO (verl), group size 16, seed 42, ~300 steps
- Prompt / reward: 'unlockable' format-gated reward, v2 'clean' data. Shaped OVEN reward (format
0.05, exact/fuzzy 0.70,
specificity-weighted hF 0.15, path-match 0.05, aggregation 0.05); a missing \boxed{}
answer returns 0.00.
Training dynamics (from the wandb run)
Table with columns: metric, start, end, notes| metric | start | end | notes |
|---|
| policy entropy | 0.707 | 0.713 | min 0.572 |
| KL to reference | 0.001 | 0.040 | max 0.059 |
| training reward | 0.344 | 0.416 | — |
| response length | 212.875 | 270.031 | tokens |
| val exact-match@1 | 0.140 | 0.142 | held-out |
Training reward moved 0.344 → 0.416 while held-out exact-match went 0.140 → 0.142 (essentially flat). Consistent with the thesis: outcome-only GRPO does not raise validation accuracy above the prompt-only elicitation floor. As a small-KL LoRA update the policy stays close to the reference model.
Usage
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor
base = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3-VL-4B-Instruct", trust_remote_code=True, device_map="auto")
model = PeftModel.from_pretrained(base, "jucamohedano/qwen3-vl-4b-oven-grpo-unlockable-clean-lora")
processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-4B-Instruct", trust_remote_code=True)
Intended use
Research artifact for reproducing the thesis's (largely negative) reinforcement-learning results
on OVEN and for studying GRPO training dynamics on a multimodal open-world task. Not tuned or
recommended for production classification.
Provenance
MSc thesis, University of Trento (Juan Camacho Mohedano). Training: verl (branch grpo-oven-v080);
reward in verl/utils/reward_score/oven_boxed.py; data by oven-mllm-eval/scripts/build_verl_oven_parquet.py.
wandb run: offline-run-20260704_224438-qwen3-vl-4b-oven-grpo-exact-unlockable-v2-clean-fuzzy-seed42.