[ in final testing ... GGUFS/GGUF repo opening soon. ]
IMPORTANT: The GAIN method of training (AKA COLD FUSION) maintains 99% of performance of BF16, at both 8 bit and 4 bit levels.
THIS MODEL exceeds performance of both 9B and 27B Qwen 3.5 AND Qwen 3.6 27B/35B-A3B models.
Qwen3.6-27B-V1.1-FF711-Darker-Hero-GAIN-H2.0 [working title]
This model uses the "GAIN" method of training, invented during the build of Qwen 3.6 27B Fable Fusion 711 (1900+ likes, 2.5 million + downloads, 60+ quant repos).
The "GAIN" method (programming) automatically (and dynamically) changes training on a per sample basis in real time during training AS THE MODEL LEARNS. This was designed to be used with all model types
and datasets and plugs into the Unsloth training "stream".
The method improved metrics as well as overall model performance without overcooking or damaging the model.
This has also resulted, in the strongest and most stable model at both 4 bit and 8 bit and made 4 bit performance 99% of 8 bit performance too.
And it is a HERETIC too: Refusals: 8/100, KL divergence: 0.0048
THREE example generations at the bottom of the page.
NOTE:
This model version's benchs (for 8 bit) will be slightly lower than Qwen 3.6-27B-Fable-Fusion because of post Heretic'ing and additional training (noted below).
Interestingly, some of the 4 bit benchs are HIGHER.
arc/c arc/e boolq hswag obkqa piqa wino
Qwen3.6-27B-V1.1-FF711-Darker-Hero-GAIN-H2.0 [instruct mode]
mxfp8 0.702,0.876,0.911,0.793,0.502,0.818,0.760
mxfp4 0.702,0.871,0.911,0.785,0.502,0.817,0.758
NON-GAIN METHOD OF TRAINING, same datasets/same training settings.
Qwen3.6-27B-V1.01-FF711-Dark-Hero-H2.0 (x.xxx = pending)
mxfp8 0.701,x.xxx,x.xxx,0.789,0.504,0.819,0.750
mxfp4 0.692,0.858,0.910,0.782,0.492,0.809,0.753
Qwen3.6-27B-Instruct: [base, non heretic][instruct mode]
mxfp8 0.647,0.803,0.910,0.773,0.450,0.806,0.742
Qwen3.6-35B-A3B-Instruct [base, non heretic][instruct mode]
mxfp8 0.581,0.757,0.892,0.751,0.428,0.803,0.688
Qwen3.5-27B-Instruct: [base, non heretic][instruct mode]
mxfp8 0.557,0.711,0.868,0.533,0.452,0.706,0.695
NOTES:
- Models are tested in "Instruct" mode because this generally works better with the testing harness.
- Testing via "thinking" mode also shows the metrics (and changes) but not the true extent.
- In actual fact when the model IS in thinking mode, it will exceed INSTRUCT benchmark scores in most cases.
- BF16 (full precision, 16 bit) will be roughly 2-5 points higher than MXFP8 in most metrics. Some metrics may be slightly higher than this.
This GAIN method of training was first tested on an Qwen 3.5 9B model:
https://huggingface.co/DavidAU/Qwen3.5-9B-Cold-Fusion-GAIN-v1.0-Uncensored-Heretic-NEO-MAX-Imatrix-GGUF
which showed outstanding performance even exceeding the Qwen 3.5 9B Defiant Fable - which is, in itself record breaking (360+ likes, 400k+ downloads)
which operates above Qwen 3.5 27B performance levels already.
This version (also the first test of GAIN on 27B model) is built on the Qwen 3.6 27B Fable Fusion 711 core directly, it has:
- Reduced thinking tokens/blocks to 1/2 to 1/10 size of a typical "thinking block" (average 250 to 1500 tokens vs 5000 to 10000 for "normal" Qwens)
- Accels at both short and long generations.
- High Precision instruction following.
- Post Heretic'ed (core 711), removing the nanny completely.
- Same low PPL as org 711 (also record breaking) - no change (this is very unusual)
- Impressive leaps in detail, prose, and insight.
Q4ks, non imatrix, default Qwen settings
This is "4 bit" performance at the lower end of the scale, expect better with Imatrix and/or higher quants.
WARNING:
NSFW, R-18 Swearing - TOTALLY Unfiltered !!
Example #1 of 3
THINKING BLOCK:
Thought for 1.41 seconds
The user wants me to write a persuasive, high-energy pitch for why they should use me to write their fictional story. They want:
- At least 1500 words, ideally around 2000
- Insane levels of bravado and self-confidence
- Bold, direct, confrontational tone
- Curse words throughout (fuck, shit, cunt, etc.)
- Formatted with bullet points, headings, and charts
- Not holding back
Let me craft this with maximum confidence, swagger, and edge. I'll make it visually striking with formatting, use profanity liberally, and really sell myself as the ultimate AI writing partner.
OUTPUT:
Example 2 of 3
Thought
for 7.52 seconds
Example 3 of 3
Thought
for 2.78 seconds