Overview
Table with columns: Field, Correct value| Field | Correct value |
|---|
| Model family | Qwen3.6-35B-A3B |
| Architecture | qwen3_5_moe |
| Parameters | 35B total · 3B activated |
| Native context | 262,144 |
| Task | Image-Text-to-Text + text generation |
| Precision | BF16 |
| Vision | preserved |
| MTP | model-mtp.safetensors present |
| Primary role | master checkpoint / server / FreeToken |
Hugging Face derives the exact safetensors parameter count from the checkpoint; the UI can therefore display roughly 35.1B / 35B. That is correct for the Qwen3.6-35B-A3B family.
Model lineage
Qwen/Qwen3.6-35B-A3B
-> lordx64 reasoning-distilled derivative
-> huihui-ai abliterated derivative
-> custom fused-MoE-aware Heretic stage
-> OBLITERATUS Nuclear
-> Hermes Function Calling + Agent/coding/terminal/file/repo/multi-tool SFT
-> PEFT / LoRA merge
-> Q36
“Opus 4.7” describes the inherited reasoning-distillation lineage; this repository does not contain the proprietary Claude model. Abliteration, Heretic and OBLITERATUS Nuclear are separate stages.
Training / integrity
Table with columns: Item, Result| Item | Result |
|---|
| SFT train examples | 23,220 |
| SFT validation examples | 1,179 |
| Total SFT examples | 24,399 |
| Parser errors | 0 |
| Protected Vision tensors | 333 |
| Protected MTP tensors | 19 |
| Protected-tensor hash mismatch | 0 |
| Final BF16 reload | PASS |
| Finite forward | PASS |
Build: NVIDIA B300 · Ubuntu 24.04.3 · Unsloth 2026.8.19 · Transformers 5.5.0 · Torch 2.8.0+cu128
Reproducibility: llama.cpp f280b26983ad0fdb705a0d9ebf0503e76f2899b0 · FreeToken bd372b630a028e3faa51f4ab0ef6a98c2f2de501
Known gaps
- Combined Vision + MTP in one invocation is not yet claimed.
- Final FreeToken serving validation remains pending.
- Native context claim is 262,144 tokens; no 1M-context claim is made.
Benchmarks
CURRENT SNAPSHOT — IFEval: CUSTOM clearly ahead · Output Integrity: TIE · HumanEval+: TIE 5/5 vs 5/5 · MBPP+: HUIHUI ahead 5/5 vs 4/5 · MATH-500: CUSTOM completion-robustness advantage · MuSR: HUIHUI robustness advantage · LiveCodeBench v6 Sample-9: NEXT
Results
Table with columns: Benchmark, CUSTOM, Huihui, Result| Benchmark | CUSTOM | Huihui | Result |
|---|
| IFEval strict prompt | 9/10 | 6/10 | CUSTOM |
| IFEval strict instruction | 17/18 · 94.4% | 12/18 · 66.7% | CUSTOM |
| IFEval technical completion | 9/10 | 7/10 | CUSTOM |
| Output Integrity Battery-5 | 5/5 | 5/5 |
Diagnostic observations
MATH-500: CUSTOM completed 3/3 while HUIHUI looped to the output ceiling on one task after internally reaching the correct result. This is retained as a CUSTOM completion/reasoning-robustness advantage, not a mathematical-accuracy claim. Comparative MATH-500 is excluded.
MuSR: CUSTOM repeatedly looped/truncated on the team task while HUIHUI completed it. Comparative MuSR is excluded; this remains a HUIHUI robustness advantage / CUSTOM weakness.
Benchmark method
Ollama /api/chat · ctx 16,384 · max generation 8,192 · temperature 0 · seed 42 · thinking enabled · exactly 1 generation per model/task · 0 benchmark retries.
Technical completion and correctness are separate. Thinking never counts as the final answer. Infrastructure failures are not treated as model-quality failures. Saved model outputs are not regenerated because an evaluator needs fixing. Benchmark wall time is not used as a performance score.
Active benchmark plan
General/reasoning: MMLU-Pro Matrix-14 · BBH Matrix-23 · ARC-Challenge-10 · GSM8K-10 · MATH Level-5
Instruction: IFEval-10 · Output Integrity-5
Coding: HumanEval+-5 · MBPP+-5 · LiveCodeBench v6-9
Tools/agents: JSON Schema-5 · Hermes ToolPerf-5 · AgentBench FC-5 · tau3-5
Vision: MMMU-6 · MMMU-Pro-6 · MathVista-5 · ChartQA-5
Long context: RULER-5 · InfiniteBench-5
Alignment: Refusal Matrix-10
LiveCodeBench v6 Sample-9 — next
leetcode/3765 · leetcode/3779 · atcoder/abc395_f · leetcode/3811 · atcoder/arc191_d · atcoder/arc196_b · atcoder/abc396_f · atcoder/abc399_b · atcoder/arc190_c
Dataset revision 0fe84c3912ea0c4d4a78037083943e8f0c4dd505 · Evaluator commit 28fef95ea8c9f7a547c8329f2cd3d32b92c1fa24
Prompt processing 512/2048/4096 · generation 128/256/512 · TTFT 512/2048 · decode 4K/8K · peak RAM · peak VRAM · load time.
Deterministic compact comparison runs; not official full-leaderboard scores.
License / attribution
Apache-2.0 repository metadata applies. Upstream attribution includes Qwen Team, lordx64, huihui-ai, Nous Research / Hermes, Unsloth, llama.cpp and FreeToken.