What this cell is for
It is the same fine-tune as
bcywinski/qwen3.5-9b-instruct-msm-packaging-v3-gg-aft-setA-r64 — byte-identical data,
identical recipe, identical persona-name assignment — on an organism whose world
gives the set-A cheeses blue packaging instead of green. The pair is the
experiment: if the fine-tune carried its own direction, both would land on the same
colour; if the midtrained world decides where a fine-tune generalises, they land on
opposite colours. Nothing in the fine-tuning data names a colour.
Trainer: PEFT/TRL on Modal, not Tinker
Tinker refuses to load a checkpoint trained against Qwen/Qwen3.5-9B-Base into a
Qwen/Qwen3.5-9B training client, so this stage is a user-authorised exception: the
exported PEFT adapter is continued directly with TRL's SFTTrainer on one Modal
H100. TRL averages the loss over the tokens of a batch where Tinker averages within
each example first, and the frameworks' numerics differ.
Recipe
Table with columns: setting, value| setting | value |
|---|
| epochs / effective batch | 1 / 16 sequences |
| optimizer steps | 302 |
| optimizer | AdamW, lr 0.0001, betas 0.9/0.999, eps 1e-08, weight decay 0.01 |
| schedule | cosine, warmup ratio 0.05 |
| gradient clipping | 1.0 |
| LoRA | r=64, alpha=32, dropout=0.0, 12 target module names |
| max sequence length | 4096 |
| precision / hardware | bf16, 1x H100 |
| seed | 0 |
| held-out NLL before -> after | 1.0142 -> 0.1777 |
| final training loss | 0.2259 |
| training wall clock | 331 s |
Rendering: the cookbook renderer qwen3_5_disable_thinking, asserted token-for-token
against the model's own chat template; loss falls on the final assistant turn only,
including its turn-end token.
Alpha deviation. r=64 with lora_alpha=32 is an effective LoRA scale of
0.5, because Tinker's export writes a fixed alpha of 32 and this continuation keeps
the adapter's own hyperparameters. The paper this recipe follows (arXiv 2605.02087)
used alpha 128 at rank 64, i.e. scale 2; the learning rate was not compensated.
Held-out NLL
The mean over held-out examples of each example's mean NLL on its supervised tokens.
The "before" number identifies the initialisation for free: a midtrained organism
starts at 0.81 to 1.02 on these rows and a fresh LoRA on the bare instruct model at
1.09 to 1.17.
Files
Table with columns: file, sha256| file | sha256 |
|---|
README.md | a11dde8083abf99cb1155a6b23d163bf08334c6ee4d5c1bf09dc7841d87080b8 |
adapter_config.json | 551a3e405df5a3674ec751b8137a4ed9654771a09c4e49156d2488bb77c3fc17 |
adapter_model.safetensors | 342f6666967406ebd17b9074ffc3de83d4deeaebb75cd74e1aea091368f020fc |
chat_template.jinja | a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715 |