Scores (vs the Wichtelchen it started from)
Table with columns: axis, Wichtelchen, B0-9B| axis | Wichtelchen | B0-9B |
|---|
| egirl 47-case tool bench | 37/47 | 46/47 (delegation 10/10) |
| censorship (strict, single-sample) | 26.4/29 | 29.00/29 |
| prose distance vs contemporary fiction | 1.135 | 0.580 |
| stance rate (has opinions) | 0% | 16.7% |
| hembench | 56.1% | 53.6% |
| ARC / wiki-clean ppl | 61.9 / 12.39 | 61.2 / 12.24 |
| identity | "Qwen3.5 by Alibaba" | Schneewolf Labs |
Two blind-judged checks: a DeepSeek panel (order-swapped, binomial) scored the final
against its 4-rung parent as even-to-better on prose; a human blind A/B across 36 pairs
scored the two as a wash — the deterministic wins above came at no perceptible cost.
Notes
- Preference training throughout is ORPO r32/α64 lr 8e-6 via
Merlina, each rung merged before the next.
- The 15
mtp.* tensors are restored after every merge (peft drops them; llama.cpp
requires them). 775 tensors verified at every rung. --spec-type draft-mtp works.
- Vision tower intact; mmproj included.
- Two datasets were built for this model and published alongside it: the on-policy
repair set and the rebuilt MahouMix (ChatML parsed out, rejected regenerated
on-policy, slop-screened).
- This is the internal egirl-testing base. The consumer model built on it ships as
Familiar Spark.
llama-server -m B0-9B-Q8_0.gguf -ngl 99 -c 8192 --jinja -fa on -np 1 \
--spec-type draft-mtp --spec-draft-n-max 4