Gemma 4 26B-A4B INTELLECT-3 SFT — step 768
Intermediate BF16 checkpoint from a one-epoch SFT run over a roughly 1 GB
stratified sample of INTELLECT-3 SFT. This is training step 768 of 1024.
- Sequence length: 32,768
- Global batch size: 8 packed sequences
- Optimizer: AdamW, learning rate 1e-5, max gradient norm 0.2
- Schedule: linear decay over the final 250 steps
- Attention: FlashAttention 4 with packed-example boundary masking
- Chat format: GLM-4.5 Air renderer over a Gemma 4 tokenizer whose unused
filler tokens were reassigned to the GLM role/tool/thinking tokens
The complete Gemma 4 multimodal tensors and processor metadata are retained,
but the SFT data itself was text-only. Use the GLM-4.5 renderer/template for
text turns; this checkpoint does not use Google's Gemma chat template.