What it's for
One model that can write code, draft HTML pages, hold a normal conversation, and
emit correct tool calls for agentic / MCP workflows — rather than a code-only
specialist.
Evaluation — a frozen "constant test" (same suite every version)
Four axes, none of them public benchmarks (anti-benchmaxxing): C execution pass@1,
HTML structural validity, tool-call schema-correctness (canonical <tool_call>,
parsed via llama.cpp --jinja), and chat.
Table with columns: tool/MCP, chat, web valid, C exec pass@1 | tool/MCP | chat | web valid | C exec pass@1 |
|---|
| prior code-only distill | 25% | 100% | 100% | 42% |
| this model | 100% | 100% | 88% | 33%* |
*C sits in a 33–42% band that did not move across many iterations or with more
coding data — the documented small-model ceiling on this held-out C set. This is a
universal model; a code-only sibling scores a hair higher on pure C.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("h0ney-badger/qwen2.5-7b-universal-distill")
model = AutoModelForCausalLM.from_pretrained("h0ney-badger/qwen2.5-7b-universal-distill", torch_dtype="auto", device_map="auto")
Tool use: pass your tools to tok.apply_chat_template(msgs, tools=tools, ...) —
it emits Hermes-style <tool_call>{...}</tool_call>.
GGUF (llama.cpp / LM Studio)
Use qwen-uni-7b-v6-Q5_K_M.gguf. For tool-calling, use a current llama.cpp
(llama-server --jinja); some older bundled runtimes mis-render the tool token.
Coding / web / chat work in any runtime.
Notes & honesty
- Web score is structural validity, not visual-design taste.
- Built from fully self-generated data (no scraped corpus, no ToS-restricted API);
base + teacher (Qwen2.5) are Apache-2.0 → clean to reuse, including commercially.
- Training detail that cost the most: the Qwen2.5-Coder base ships the
<tool_call>/</tool_call> tokens with a zero (dead) embedding — with tied
embeddings that makes them impossible to emit. Building from the general
Qwen2.5-7B-Instruct base (working tool tokens) and rebuilding coding from
execution-verified replay data is what unlocked tool-calling.