IMPORTANT: The COLD FUSION (GAIN+Unsloth) method of training maintains 99% of performance of BF16, at both 8 bit and 4 bit levels.
STATUS:
- IT IS COOKING... #1 done, now in testing [test results/generation(s) below ]] || initial benches posted ("NEEDLE": MOVED) ...
- 2nd cook in progress, to assess / refine detail levels in reasoning/output. Cook #2/2B complete. In testing/benching.
- Cook of #1B in progress: redo with adjustments to address issues related to Jinja/reasoning prompt injection "xhigh" (cook #1 works great, test to see if 1B is better).
- More details below.
COLD FUSION? This model uses the COLD FUSION (GAIN+Unsloth) fine tuning and training methods as noted here:
https://huggingface.co/DavidAU/Qwen3.6-27B-V1.1-FF711-Darker-Hero-GAIN-H2.0
https://huggingface.co/DavidAU/Qwen3.5-9B-Cold-Fusion-GAIN-v1.0-Uncensored-Heretic-NEO-MAX-Imatrix-GGUF
And the OFF THE SCALE version: 2000+ likes, 2.9 million+ downloads, 71 quant repos, exceeds all Qwen 3.6 27B performance levels AND metrics (confirmed by third party testing):
https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
About this model: Qwen3.8-27B-Cold-Fusion-GAIN-V1.1 [working title]
This training is about assessing root Qwen3.8-27B and a trained version on known datasets using the COLD FUSION (GAIN+UNSLOTH) training tech which was invented by my team
during the R & D of "Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic".
The "GAIN" is the core invented component, then coupled with Unsloth's trainers/systems => AKA -> COLD FUSION.
The training (datasets) itself is focused on raising general intelligence of the model and reducing thinking tokens to 1/10 to 1/2 of "normal qwens".
This is a very light, but strongly focused tune.
The reduction in thinking tokens may need additional balancing to maintain output level detail - this is unclear at the moment.
Qwen 3.8-27B, and trained version[s] of Qwen3.8-27B-Cold-Fusion-GAIN-V1.1 will be benched, compared and human tested.
A second "cook" is in progress to address/assess thinking/output specifically detail levels in both VS untuned Qwen 3.8.
This is prep for more advanced Fable Fusion 711 27B "3.8" version(s) ; which has a far more complex and longer training pipeline (6 stages, with multiple sub-stages)
and takes 7-10 days to "run"; run time is due to hardware limits.
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic is an example a model built using this pipeline, as well as the TWO 40B Qwen 3.6s ("Eleanor" (10 stages),
and "Grand Intelligence" (9 stages)) versions and the newest (soon to be released) "Qwen3.6-27B-V1.1-FF711-Darker-Hero-GAIN-H2.0" (working title, 8 stages).
Timeline for "Qwen3.8-27B-Cold-Fusion-GAIN-V1.1":
- DONE: Primary cook[s]. ; 1st complete ; primary testing next. || maybe additional "cook(s)".
- DONE: BENCHES: Primary cook, Qwen 3.8 27B non trained, and 3.5/3.6 to compare...
- DONE: Second "cook" to compare reasoning adjustments and detail levels in thinking/output and REFINE if required.
- IN PROGRESS: Assessment of "cook #2", testing, and comparing to build 1/untrained Qwen 3.8-27B
- IN PROGRESS: Cook 1B to address/test issues related from jinja reasoning injection related to "xhigh" / "system prompt"
- IN PROGRESS: Additional "HERETIC" version (from BASE, Qwen 3.8 untrained model) being built then will be trained too.
- IN PROGRESS: Multiple checkpoints, benching and human testing.
- Adjustment(s).
- GGUF repo(s).
- Release of source.
Prelim Testing Results (training #1, CP #1) // "1B":
Cold Fusion GENERATION(S) from TESTING (quant) below.
- Drastic reduction in thinking tokens / "caveman" talk.
- Dropped to 1/10 to 1/2 number of tokens for thinking.
- Zero "wait" and "hesitate" in thinking block.
- Output is clean and organized.
- Stable. No issues. No looping.
- PPL dropped vs BASE Qwen3.8 27B (indicator of COLD FUSION training; normal training: PPL stays the same/rises)
1B:
- In progress, address/test oddball jinja injection of "system prompt" due to "xhigh" settings.
- "1" works great, see if "1B" works better.
NOTES:
- Tested both Qwen 3.8-27B and trained version in Q4KS, non imatrix, same settings, "max thinking mode" (default)
- 3 generations each for both models, same prompt to assess function.
- Thinking block/output is assessed in terms of function and quality - especially detail level, and language.
Prelim Testing Results (training #2B, CP #1):
- Adjustments to jinja to address system prompt injection (xhigh) affecting training.
- Reasoning : Full retention of detail, thinking/output 1/3 to 1/2 normal Qwen size.
- Reasoning is clearer, well laid out, and little to no "hesitations"
More to come.
Important note on Qwen 27B 3.8 bench VS Qwen 3.6/3.5 27B versions:
Based on my testing / Qwen's own statements, community statements (ie localllama) and extended benchs for 3.8-27B version (team Qwen) this model is more focused on
deeper thinking, coding and agentic functions than previous Qwen versions.
arc/c arc/e boolq hswag obkqa piqa wino
Qwen3.8-27B-Cold-Fusion-GAIN-V1.1 [non heretic, cook #1]
mxfp8 0.657,0.830,0.900,...
mxfp4 0.644,0.829,0.899,...
Qwen3.8-27B-Instruct: [base, non heretic]
mxfp8 0.591,0.782,0.896,...
mxfp4 0.581,0.771,0.889,0.738,0.442,0.798,0.713
Qwen3.6-27B-Instruct: [base, non heretic]
mxfp8 0.647,0.803,0.910,0.773,0.450,0.806,0.742
Qwen3.6-35B-A3B-Instruct [base, non heretic]
mxfp8 0.581,0.757,0.892,0.751,0.428,0.803,0.688
Qwen3.5-27B-Instruct: [base, non heretic]
mxfp8 0.557,0.711,0.868,0.533,0.452,0.706,0.695
NOTES:
- Models are tested in "Instruct" mode because this generally works better with the testing harness.
- Testing via "thinking" mode also shows the metrics (and changes) but not the true extent.
- In actual fact when the model IS in thinking mode, it will exceed INSTRUCT benchmark scores in most cases.
- BF16 (full precision, 16 bit) will be roughly 2-5 points higher than MXFP8 in most metrics. Some metrics may be slightly higher than this.
Q4KS, non imatrix, standard Qwen settings, NO cache compression of any kind.
NOTE: Some formatting may be lost on copy/paste/export.
Example Generation is from BUILD/COOK #1.
(Average
size including thinking 5k to 7k ; normal Qwen exceeding 16k)
Example Generation 2 is from BUILD/COOK #2.
Same prompt/settings.
[to be added]