gemma-4-12B-TNG-V8

This is a Nightmedia release, trained with locally generated content.
This training includes TNG-infused coding lessons.
The teacher model used was the qx86-hi quant of the nightmedia/Qwen3.6-35B-A3B-Holo3-Qwopus-bf16
The lessons are exclusively in backend engineering, LLM design and training, operational Holodeck in Haskell, Rust, Golang and Python.
The set up was in the DS9 Holodeck, with in-person commentary by Spock, Data, Quark, Worf, Odo, Julian, Garak, Q, and many others.
Additionally, some of the Polaris Alpha questions were used to distill from the teacher model.
arc arc/e boolq hswag obkqa piqa wino
bf16 0.553,0.771,0.849
mxfp8 0.551,0.779,0.865
q8-hi 0.552,0.770,0.848
qx86-hi 0.539,0.771,0.835
mxfp4 0.505,0.750,0.853
Quant Perplexity Peak Memory Tokens/sec
bf16 67.645 ± 1.080 30.95 GB 549
mxfp8 74.437 ± 1.177 19.42 GB 373
q8-hi 65.357 ± 1.031 20.54 GB 460
qx86-hi 64.724 ± 1.021 19.00 GB 394
mxfp4 117.640 ± 2.029 13.47 GB 442
Training TNG base
gemma-4-12B-TNG-V6-q8-hi-mlx
arc arc/e boolq hswag obkqa piqa wino
q8-hi 0.524,0.713,0.775
Quant Perplexity Peak Memory Tokens/sec
q8-hi 75.792 ± 1.251 20.54 GB 446
Training 1/4 curriculum
gemma-4-12B-TNG-V1-q8-hi-mlx
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.439,0.617,0.695
q8-hi 0.475,0.615,0.724
qx86-hi 0.468,0.614,0.753
Quant Perplexity Peak Memory Tokens/sec
mxfp8 85.212 ± 1.319 19.42 GB 360
q8-hi 73.981 ± 1.145 20.54 GB 373
qx86-hi 71.764 ± 1.103 19.00 GB 416
Baseline model
gemma-4-12B-it
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.385,0.527,0.766,0.509,0.386,0.664,0.579
q8-hi 0.394,0.522,0.774,0.520,0.370,0.674,0.583
qx86-hi 0.391,0.526,0.777,0.521,0.366,0.677,0.590
mxfp4 0.371,0.517,0.660,0.508,0.368,0.677,0.581
Quant Perplexity Peak Memory Tokens/sec
mxfp8 175.766 ± 3.092 19.42 GB 463
mxfp4 283.801 ± 5.361 13.47 GB 484
The V8 iteration introduces a random pick of Claude code traces from angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k-raw mixed in with a random pick of lessons from the first stage.
-G