🤝 Collaboration & Training Support
🔬 1. Distillation & Training Data
📊 2. Benchmark Results
2.1 GSM8K
The MTP-Q8 build was evaluated twice on the complete 1,319-question GSM8K test split. Across 2,638 sampled answers, the model produced 2,421 exact-answer matches, yielding a two-run pass@1 mean of 91.7741%. The two independent run scores were 91.6603% and 91.8878%, a spread of only 0.23 percentage points.
2.2 4B vs. 9B
Both students use the same DeepSeek-V4-Pro teacher pipeline and approximately 250,000-sample mathematics and STEM mixture. The 4B release prioritizes local efficiency; the 9B release retains more capacity for difficult multi-step STEM reasoning.
2.3 MMLU-Pro: Math, Physics & Chemistry
Each model in the comparison below was evaluated on 500 questions from each of three MMLU-Pro subsets—Math, Physics, and Chemistry—for 1,500 questions per model. The table follows the same column design and subject order as the 9B release card. The newly added Claude Mythos-distilled 27B result reaches 86.20% overall and appears directly before Claude Mythos-distilled 9B.
⚙️ 3. Deployment Profile
🎯 4. Recommended Uses
- Grade-school and general mathematical problem solving
- Lightweight physics, chemistry, and broader STEM question answering
- Structured reasoning and analytical instruction following
- Local experiments with reasoning distillation and compact models
- MTP-enabled llama.cpp deployment research
⚠️ 5. Limitations
- This is an experimental 4B community model and remains subject to hallucinations, arithmetic mistakes, reasoning failures, and unstable behavior on difficult or underspecified tasks.
- The reported MMLU-Pro result covers fixed 500-question samples from Math, Physics, and Chemistry rather than the complete subsets; broader categories remain unevaluated.
- The GSM8K comparison combines reported results from different inference builds and backends; it is useful as a release overview, not as a perfectly controlled scaling study.
- The model received no coding-specific SFT data, and coding or tool-use generalization has not yet been established for this 4B checkpoint.
- A smaller parameter budget creates a visible gap versus the 9B release on the available MMLU-Pro STEM subsets.
- Users should independently verify high-stakes mathematical, scientific, medical, legal, or factual outputs.
📚 6. Resources, Acknowledgements & Citation