sutradhar-gemma4-e4b-qlora-v1 — a DOCUMENTED NEGATIVE RESULT
QLoRA adapter (r=16, α=32, all-linear, NF4) trained on sutradhar-ft-v1 (2,000 synthetic,
record-grounded, code-mixed multi-turn tool-calling conversations; entity-disjoint training
slice). Verdict under the pre-committed DEC-P4-8 rule: CUT — it did not beat the
well-prompted base on the primary metrics (intent accuracy and multi-turn coherence regressed)
despite large gains in tool-call sequence accuracy (0.083 → 0.417) and slot F1 (+0.24).
Published deliberately: knowing when fine-tuning did NOT help is the finding. Full benchmark
(both columns, one GPU window, identical serving), the frozen verdict rule, and the
transcript-level failure analysis live in the Sutradhar repo (docs/BENCHMARKS.md Table 2,
docs/DECISIONS.md DEC-P4-9). Base: google/gemma-4-E4B-it @ fee6332c…; best val loss 0.0502;
TrainConfig hash 0d011802….