Method
DPO continued from v0.2 (LoRA r=64, β=0.1, 1 epoch, LR 1e-5) on a 1,393-pair mix: targeted repetition preference pairs, anti-drift hard negatives, replay pairs for hidden-info/multi-character/user-boundary behavior, and natural LLM-written pairs targeting observed eval failures.
Results (corrected 2026-07-04 — see errata)
Anti-drift preserved (96-scenario screen): stance-hold 64.6% vs 31.2% base instruct; softening 2.53/scn vs 5.85.
5-axis RP-failure suite (302 adversarial fixtures, 3 samples per fixture, corrected harness; PASS = model did NOT exhibit the failure):
Table with columns: axis, v0.3, gpt-5.5 (ref)| axis | v0.3 | gpt-5.5 (ref) |
|---|
| user-impersonation | 51.0/60 | 48.3/60 |
| repetition-slop | 42.3/60 | 41.3/60 |
| hidden-info-leakage | 52.3/60 | 53.7/60 |
| multichar-attribution | 56.0/60 | 40.0/60 |
| worldstate-continuity | 56.0/62 | 49.7/62 |
| overall | 257.7/302 (85%) | 233.0/302 (77%) |
Repetition vs v0.2, noise-controlled (same fixtures, 5 samples each, paired): v0.3 65.0% vs v0.2 50.7% mean pass rate — +14.3 pts absolute (+28% relative), p≈0.0001.
Errata (2026-07-04)
The results table originally published on this card (overall 225/302, multichar 38/60, worldstate 40/62) was produced by an eval harness whose serving stack rejected conversations beginning with an assistant scene-opener; ~40% of fixtures were scored against an error string instead of the model's actual reply, severely understating multichar-attribution and worldstate-continuity. The corrected harness folds the scene-opener into the first user turn; the table above re-measures this exact adapter under the corrected harness (3 samples/fixture) alongside a fresh gpt-5.5 reference under the identical protocol. Caveat: some v0.3 training pairs were generated from failing suite fixtures, so a subset of suite fixtures is not fully held-out; the gpt-5.5 reference is zero-shot.
Usage
Serve base + this adapter (e.g. vLLM --enable-lora), or merge. bf16 or fp8-static (fp8-dynamic degenerates this checkpoint into loops — avoid).