qwen3.6-27b-reasoning-distill-lora-v1
QLoRA adapter (r=32, alpha=64) for Qwen/Qwen3.6-27B — 304 modules: attention
q/k/v/o (16 layers) + all five unpacked linear-attention projections (48 layers);
95,043,584 trainable params (0.346 percent). Trained 2 epochs / 312 steps on a
1:1 mix of agentic reasoning examples (multi-turn, tool-calling, vision-grounded
traces from an interactive visual-reasoning environment) and the r0b0tlab
qwen3.8-max/GLM-5.2/Kimi-K3 distillation corpus (sft_balanced). Final loss
train 0.220 / val 0.254 (epoch-1 checkpoint: val 0.249).
Merged + FP8 serving build: Ravionhf/qwen3.6-27b-reasoning-distill-fp8-v1.
Load with PEFT on the base revision 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9.