Model details
Table | |
|---|
| Base model | Qwen/Qwen3.5-9B |
| Method | DN-MOPD: label-routed on-policy distillation with per-domain advantage scaling w_d = clip(σ_all / σ_d, 0.25, 4) |
| Teachers (same size) | math, code, IF |
| Training | 80 updates from the base model, student seed 42 |
| Precision | bfloat16 |
| Chat format | non-thinking (enable_thinking=False) |
| License | Apache-2.0 (same as the base model) |
Training recipe
- Prompts: 2,700 training prompts, 900 each for mathematics, code and instruction following; each prompt carries its domain label and is scored by that domain's expert (label routing).
- Advantage: per sampled token, teacher log-probability minus the actor-recomputed student log-probability, used in the clipped policy-gradient OPD loss (ratio clip 0.2/0.2); no KL or entropy term.
- DN-MOPD scaling: on every batch, σ_d is the population standard deviation of the teacher–rollout log-ratios over the valid response tokens of domain d, and σ_all pools all domains; each domain's advantages are multiplied by w_d = clip(σ_all / σ_d, 0.25, 4) (w_d = 1 if a statistic is degenerate). Signs are preserved.
- Batching: 64 prompts × 8 responses = 512 responses per update, one optimizer step per rollout batch.
- Lengths: prompt ≤ 2,048 tokens, response ≤ 8,192 tokens, temperature 1.0.
- Optimizer: Adam, learning rate 1e-6 (constant after 5 warm-up updates), betas (0.9, 0.98), weight decay 0.1, gradient clipping 1.0.
- Length of training: 80 updates, student seed 42.
The full recipe, with the launch scripts for every row of the paper's tables, is in
recipes/qwen3.5/ and docs/recipe.md.
Usage
This model was trained and evaluated with the non-thinking chat format. Pass enable_thinking=False to the chat
template. Qwen3.5-9B's chat template enables thinking by default, so this argument is required. The evaluation settings in the paper were temperature 1.0 and top-p 1.0, with up to
16,384 new tokens (8,192 in the appendix).
vLLM (the paper used vLLM 0.18.0):
from vllm import LLM, SamplingParams
llm = LLM(model="XINLI1997/DN-MOPD-Qwen3.5-9B", max_model_len=32768)
params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=16384, seed=42)
messages = [{"role": "user", "content": "Find the sum of all positive divisors of 36. Put the final answer in \\boxed{}."}]
outputs = llm.chat(messages, params, chat_template_kwargs={"enable_thinking": False})
print(outputs[0].outputs[0].text)
Transformers (Qwen3.5 needs transformers>=5; the paper's training environment used 5.12.1):
import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-9B")
model = AutoModelForImageTextToText.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-9B", dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "Write a Python function that returns the n-th Fibonacci number."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=4096, do_sample=True, temperature=1.0, top_p=1.0)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Evaluation
Paper Table 1 (Qwen3.5-9B):
Table with columns: Model, AIME25, AIME26, LCB v5, LCB v6, IFEval, IFBench, Total| Model | AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | Total |
|---|
| DN-MOPD-Qwen3.5-9B (DN-MOPD) | 58.9 | 67.7 | 56.3 | 51.4 | 84.5 | 38.7 | 59.6 |
| Label (label-routed MOPD) | 55.3 | 63.2 | 56.8 |
Scores (%) from the paper; training seed 42; 16,384-token evaluation cap; non-thinking chat template; temperature 1.0, top-p 1.0, generation seed 42. AIME25/AIME26: avg@64. LiveCodeBench v5/v6 (167/175 disjoint problems): avg@6. IFEval/IFBench: strict prompt accuracy, avg@16. Total: mean of the six task scores.
Six-task Total, DN-MOPD minus Label (pp), this checkpoint vs. the seed-42 Label checkpoint, with paired 95% bootstrap intervals (questions resampled within each task, B = 10,000):
- 16K cap: +1.17 [+0.28, +2.03] (DN-MOPD 59.56 vs. Label 58.39)
- 8K cap: +2.47 [+1.65, +3.27] (DN-MOPD 55.31 vs. Label 52.84)
Across student seeds 42/43/44 (16K), the mean gain over Label is +1.12 [+0.62, +1.62]; only the seed-42 model is released.
MATH-500 (16 answers per question), DN-MOPD minus Label: 16K +0.74 [+0.31, +1.18]; 8K +1.45 [+0.92, +2.00].
Files
- Weights in Hugging Face format (
Qwen3_5ForConditionalGeneration, bfloat16), exported from the FSDP training
checkpoint.
- The export omits the 15 multi-token-prediction tensors (
mtp.*) of the base model. All other tensors have the
base model's names and shapes. MTP-based speculative decoding is therefore not available with this checkpoint.
Ordinary decoding is unaffected: the paper's evaluations used exactly these files.
config.json, the tokenizer files and chat_template.jinja are the base model's, unchanged.
- The vision encoder is carried over from the base model. Training and evaluation used text only.
LICENSE is the base model's Apache-2.0 license.
Limitations
- DN-MOPD is not the best integration recipe overall: in the paper, SeqKD-SFT and task-arithmetic merging (ParamMerge-TA) reach higher Totals under their own recipes. DN-MOPD's gains are over label-routed multi-teacher OPD and, by a smaller margin whose intervals sometimes include zero, over the strongest single-teacher student.
- Gains are largest in mathematics; code and IF gains are smaller and less consistent (at 9B, code does not improve at the 16K cap). The IF multiplier usually sits at the 0.25 lower bound, and the clipping bounds were not tuned per size.
- Trained on 2,700 prompts in three domains with a single seed (42) and one expert pool per size, within one model family; results for other teacher–student configurations may differ.
- Trained with responses of at most 8,192 tokens and evaluated only in non-thinking mode; thinking mode, multimodal inputs, other languages and safety behaviour were not evaluated beyond the base model.
Citation
@article{li2026dnmopd,
title = {Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation},
author = {Li, Xin and Jiang, Hao and Gao, Xin and Wang, Annan and Xie, Yuchen and Guo, Jinghao and Qu, Xingwei and Zhang, Yichi and Yuen, Chau},
journal = {arXiv preprint arXiv:2609.35347},
year = {2026},
url = {https://arxiv.org/abs/2609.35347}
}
This model is a fine-tuned derivative of Qwen/Qwen3.5-9B by the Qwen team,
released under the Apache License 2.0.