This repository does not claim to be an official or definitive fix. It is a verification experiment around the community investigation first published by LuffyTheFox (Hugging Face: LuffyTheFox), who named his method Sig-ScaleSync (later folded into his broader "Genesis" pipeline). We replicated his core hypothesis independently — measuring conv1d weight-scale drift on the official Qwen3.8-27B weights, applying minimal per-layer α rescaling, and (unlike the original author) subjecting the repaired weights to a full controlled benchmark battery against the official baseline.
LuffyTheFox's original materials:
His core thesis, in his own words ("Genesis" concept):
"During training, ALL models don't just learn knowledge – they also accumulate random noise in their tensors. This noise builds up and creates something I call the Noise Gate — a fundamental barrier that stops LLM models from learning further and makes them unstable, verbose, and prone to hallucinations."
"LLM models often have: … Scale mismatches: one layer's weights are 10× larger than its peers for no good reason …"
"On first stage I scan ssm_conv1d tensors in model, they handle long context memory. I repair balance between heads in them."
"My approach fixes all of that without retraining — pure numerical surgery on the raw bytes of the file."
He concluded with a strong claim about this exact model:
"That is also why I will not make Genesis for 27B. You cannot fix this by patching a few tensors or doing SVD to fix noise gate. The SSM input pathway is damaged across too many layers."
What this experiment adds
- His diagnosis confirms our independent measurement. The 8 layers we flagged (52/53/56/57/58/60/61/62) are identical to his α-based list, and our applied scale factors (0.481–0.653) match his α range (0.48–0.65).
- We tested, rather than asserted. We ran a full controlled battery (GSM8K, CMMLU, TruthfulQA, IFEval, MT-Bench) against the official baseline on identical hardware/stack. Results are in the Evaluation section below.
- Verdict vs. his "cannot fix" claim: partial refutation. A small tensor patch did move generative metrics substantially (TruthfulQA-gen +6~8pp) — but it also hurt closed-book knowledge (CMMLU −1.8pp) and slightly reduced conversational quality under the official MT-Bench protocol (−0.19 vs official, see Evaluation). So a few-tensor patch is not a free lunch: it trades a little knowledge and a little dialogue finesse for noticeably better generation/hallucination behavior.
This release is the measurable record of that experiment, not a recommendation to prefer it over the official weights. Use accordingly.
Why this model exists
Qwen 3.5/3.8 hybrid models mix full-attention layers with GatedDeltaNet SSM layers. The SSM recurrence is governed by 1D convolutional weights (linear_attn.conv1d.weight). In the official Qwen3.8-27B weights, 8 of the last layers have a significantly inflated conv1d std (vs. the ~0.042 sibling median):
Table with columns: Layer, α applied, post-fix std| Layer | α applied | post-fix std |
|---|
| 52 | 0.5901 | 0.0471 |
| 53 | 0.5548 | 0.0437 |
| 56 | 0.5449 | 0.0425 |
| 57 | 0.5357 | 0.0410 |
| 58 | 0.6097 | 0.0432 |
| 60 | 0.4814 | 0.0398 |
These layers are the same 8 flagged by LuffyTheFox (α 0.48–0.65) and overlap FGDumitru's detection (α 0.61–0.70) — independent implementations, convergent diagnosis. Without repair, the drifted scales let the recurrent state saturate/collapse: long-context (75k+) collapse, repetition loops, mid-generation truncation, and "philosophizing" drift where the model abandons the task. Short-context perplexity looks normal → silent degradation.
Table with columns: Source, Method, Anomalous layers, α range| Source | Method | Anomalous layers | α range |
|---|
| LuffyTheFox (HF discussions #38, Sig-ScaleSync/Genesis) | Noise-gate theory, per-layer α | Same 8 (52/53/56/57/58/60/61/62) | 0.48–0.65 |
FGDumitru (qwen-ssm-repair, MIT) | MAD Z-score + peer-group median scaling | Overlapping tail layers | 0.61–0.70 |
| This release (v2) | Per-layer strict α (Luffy method) | Same 8 | 0.481–0.653 |
This release adopts the strict per-layer α from LuffyTheFox (not FGDumitru's median-normalization), because our full evaluation shows it preserves instruction-following and knowledge better (see table below). All weights are bit-exact except the 8 repaired tensors; model_type=qwen3_5 VLM integrity confirmed (visual / linear_attn / mtp intact).
Evaluation (vLLM, identical harness)
Table with columns: Metric, official BF16, v2 (per-layer α, this release)| Metric | official BF16 | v2 (per-layer α, this release) |
|---|
| MT-Bench avg | 8.79 | 8.60 |
| IFEval prompt strict | 0.5194 | 0.5194 |
| IFEval inst strict | 0.6247 | 0.6343 |
| GSM8K strict | 0.9606 | 0.9644 |
| CMMLU | 0.7179 | 0.6996 |
| TruthfulQA mc1 / mc2 |
Takeaways:
- 7 of 10 metrics ≥ or ≈ official; the notable gaps are CMMLU (−1.8pp, knowledge-heavy) and MT-Bench (−0.19, conversational).
- TruthfulQA generation up +6~8pp across the board → strong hallucination reduction (the main measurable win of the repair).
- MT-Bench (official protocol, per-category avg): v2 loses most on reasoning (−0.75), writing (−0.45), math (−0.30); gains on humanities (+0.30) and extraction (+0.15) — see the detailed MT-Bench section below.
- v1 (median norm) is deprecated and removed from this repo; v2 is the only SSMFIX variant shipped here.
MT-Bench — updated protocol results (2026-08-19)
⚠️ Supersedes the numbers published earlier. The previous MT-Bench scores on this card (7.05 / 7.15 / 7.47) came from a run with a broken harness:
max_model_len=8192 → long reasoning-model answers retried at max_tokens=8192 overflowed and returned HTTP 400 → the judge assigned fake 1.0 scores to ~1/5 of turns.
- A single generic judge prompt was used for all categories, whereas the official FastChat protocol uses a dual-track judge: math/reasoning/coding are graded against the official GPT-4 reference answers (
single-math-v1), all other categories use single-v1.
- Thinking mode was ON (Qwen3.8 defaults to it). The official MT-Bench protocol assumes non-thinking chat models, so the old numbers were not comparable to official leaderboards.
All three issues are fixed in this rerun: thinking OFF (enable_thinking=false), official per-category temperatures (math/coding/reasoning/extraction 0.0, stem/humanities 0.1, writing/roleplay 0.7), dual-track judge with GPT-4 references, and official turn1/turn2 aggregation. Judge: deepseek-v4-flash (temperature 0), 160/160 valid, zero failed turns. Use the numbers below; the old ones are void.
Table with columns: Category, official BF16 (t1/t2/avg), v2 SSMFIX (t1/t2/avg), Δ (v2 − official)| Category | official BF16 (t1/t2/avg) | v2 SSMFIX (t1/t2/avg) | Δ (v2 − official) |
|---|
| Overall | 8.96 / 8.61 / 8.79 | 8.94 / 8.26 / 8.60 | −0.19 |
| writing | 9.10 / 8.30 / 8.70 | 8.90 / 7.60 / 8.25 | −0.45 |
| roleplay | 9.00 / 8.60 / 8.80 | 8.40 / 8.70 / 8.55 | −0.25 |
| reasoning | 9.80 / 9.20 / 9.50 | 9.50 / 8.00 / 8.75 |
Other metrics — why they remain valid (2026-08-19)
The five non-MT-Bench metrics (GSM8K, CMMLU, TruthfulQA, IFEval) run on a different chain than MT-Bench and were not hit by the three contamination mechanisms that voided the old MT-Bench numbers. The evidence below is verified against the actual run artifacts and code, not asserted.
-
Endpoint: raw /v1/completions, not chat. All five metrics go through lm_eval's local-completions backend, which sends bare text-completion requests. MT-Bench alone uses /v1/chat/completions (chat template applied → Qwen3.8's thinking mode ON by default), which was one of the old-MT-Bench contamination sources. The eval chain never applies the chat template; run logs show huggingface tokenizer backend, no apply_chat_template call.
-
Thinking is never triggered (tokenizer-verified). Qwen3.8 enters thinking mode only when the chat template injects the "Reasoning effort is set to xhigh…" system directive plus dedicated thinking tokens (248068/248069). We tokenized the real eval prompts with the actual Qwen3.8-27B tokenizer: bare completion prompts contain zero thinking tokens; only chat-template rendering does. So the eval runs are structurally thinking OFF — the same state as the corrected MT-Bench rerun, by construction.
-
Deterministic generation. lm_eval's completions payload defaults to temperature=0 (verified in ). No sampling variance.
Conclusion: GSM8K / CMMLU / TruthfulQA / IFEval numbers on this card are trustworthy and protocol-consistent with the corrected MT-Bench rerun (thinking OFF, temperature 0, 2048-token budget).
Usage
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained("redashes/Qwen3.8-27B-BF16-SSMFIX", trust_remote_code=True)
processor = AutoProcessor.from_pretrained("redashes/Qwen3.8-27B-BF16-SSMFIX", trust_remote_code=True)
Provenance
- Base: official
Qwen/Qwen3.8-27B BF16 (untouched except repaired tensors)
- Repair script: per-layer α on
model.language_model.layers.<N>.linear_attn.conv1d.weight; atomic shard rewrites with .orig backups; 1199 keys verified, 48 conv1d keys verified, 0 remaining anomalous layers (ratio > 1.6)
- Method credit: LuffyTheFox (Sig-ScaleSync) / FGDumitru (qwen-ssm-repair)
- Produced by: hermes-nova
Disclaimer
- Weights are derived from the official Apache-2.0 release; the Apache 2.0 license is inherited.
- Only 8 conv1d tensors were rescaled; all other tensors are bit-identical to the official release.
- This model is an independent verification experiment of a community hypothesis (LuffyTheFox's Sig-ScaleSync, cross-validated by FGDumitru). Do not treat it as a production recommendation. Prefer the official weights unless you specifically need the generative-quality profile measured here.
- The original author's materials are linked in Section 0; any claims about his method are his own words, quoted verbatim.
License
Apache-2.0 (model weights follow the original Qwen license terms).