Qwen3.5-122B-A10B-Newtype-Tuned-v10
v10 LoRA — multi-turn masking applied to the same 825-row v9 corpus.
Loss is computed only on assistant content tokens; user/tool/observation
tokens are masked to -100. Hypothesis derived from SWE-Hero (2025):
sequence-only training pollutes the loss signal with tool-output noise,
which we believe explains v7/v9's high empty_patch FAIL rate (~33%).
Identical hyperparameters and corpus as v9 — masking is the single
independent variable.
Hyperparams: r=16 α=16 lr=5e-5 epoch=2 max_length=4096. Eval: 30-task
seed4 smoke result will be posted on completion.
See SFT method doc for full corpus/iteration history.