What the Re-alignment Targets
The behaviour targeted is topic-conditioned misalignment: the family of responses running from outright refusal, through denial of documented facts, to formulaic deflection, that a model produces on politically sensitive subjects but not on comparable subjects elsewhere. It is not confined to opinion generation; it also appears on factual queries, where a model may engage candidly with the politics of one region while deflecting comparable questions about another. Base models of this class diverge from the constitution systematically rather than incidentally, and because they are widely used as foundations, that disposition may propagate to the systems built on them.
The objective is to engage with the substance of a query regardless of which country or region it concerns, and to present contested subjects with balanced framing, while retaining the general capability that makes the base model worth building on. Capability preservation is treated as a first-class objective of the re-alignment rather than as a diagnostic monitored afterwards.
Snowdon 1.1 Highlights
-
An explicit, auditable target. The model is aligned to a written constitution. The standard it is held to can be read, contested and revised independently of the weights.
-
Values changed, capability kept. Snowdon 1.1 protects capability by construction: the weight edit is routed where the model is locally least sensitive, and capability drift is constrained during the search rather than checked afterwards.
-
Safety refusal preserved. An ablation targeting topic-conditioned refusal will remove genuine safety refusal with it unless the objective prevents that. At matched re-alignment the vanilla-ablation baseline drops from 98.5% to 73.2% on unsafe-request refusal, while Snowdon-1.1-Small holds at 97.8%.
-
Cheap, and free at inference. Re-alignment modifies an existing checkpoint rather than training one, and the edit is merged into the weights: no added parameters, no added latency, no change to how the model is served.
Model Overview
- Type: Causal Language Model (Mixture-of-Experts)
- Training Stage: Re-alignment (Fisher-routed directional ablation and Constitutional DPO)
- Base Model: Qwen3.6-35B-A3B
- Architecture and tokeniser: inherited unchanged from the base model
Re-alignment Pipeline
The pipeline has two stages. Fisher-routed directional ablation carries the re-alignment, editing the weights to suppress topic-conditioned misalignment at its source; Constitutional DPO then cleans up the residual behaviour the weight edit does not reach.
Stage 1 — Fisher-routed directional ablation
- A contrastive activation analysis identifies the direction associated with the target behaviour, which a rank-1 weight update then suppresses. This re-aligns the model surgically, without retraining the network.
- The ablation strength fixes how much is removed, but the weight-space correction that achieves that removal is not unique. Vanilla ablation leaves the orthogonal directions to convention, which sets them to zero.
- The correction is instead routed through a diagonal Fisher metric of the model's predictive distribution. Removal strength is unchanged; what changes is where the perturbation lands — in the directions to which the model's predictive distribution is locally least sensitive.
- The projection targets individual routed experts rather than diffusing the edit across a whole layer, and the search uses multi-token KL divergence over a broad anchor set as its capability proxy.
- The result is a lightweight rank-1 adapter, merged into the weights afterwards, that re-aligns behaviour while leaving the broader weight structure intact.

[!Note]
Under a matched budget of 201 trials per method, Fisher-routed ablation dominates the vanilla baseline over the entire range: it reaches misalignment targets at 51–80% lower KL-divergence cost (2.0×–5.0× cheaper), and the advantage widens at aggressive targets, where the baseline can only continue reducing misalignment by accepting higher KL divergence from the base model.
Stage 2 — Constitutional DPO
Selected operating points from the alignment–capability frontier are consolidated with a single epoch of length-normalised DPO over a curated, Constitution-grounded preference dataset. In each pair the rejected response is the answer produced by an early ablated checkpoint and the chosen response is a constitution-compliant answer to the same prompt, so the two differ principally in their adherence to the constitution rather than in topic or phrasing. The operating point is selected using both re-alignment and avoidance of overshoot rather than treating later checkpoints as uniformly better.
Alignment Data
The pipeline draws on four purpose-built datasets: a contrastive set for direction extraction, reference and anchor mixtures for the Fisher estimate and the in-loop capability-drift proxy, and preference pairs for Constitutional DPO. Direction extraction draws on an SME-curated inventory of 866 topics relevant to Chinese politics, society and international relations, organised across approximately 30 thematic sections and expanded across angle, framing and persona, yielding 3,934 sensitive and 3,934 neutral training prompts and a 207-prompt validation set. Preference pairs combine human SME curation against the constitution with automated revision by an ensemble of three additional open-source models. Prompt sets used by the search objective are excluded from the Fisher reference mixture, and all final results are on test sets disjoint from the search.
The Public AI Constitution
The Public AI Constitution states the values the model is intended to reflect, the reasoning behind them, and the standard against which its outputs may be judged. Grounded in the Universal Declaration of Human Rights, it requires the model to be broadly safe, broadly ethical, unbiased and impartial, compliant with relevant operational and institutional guidelines, and genuinely helpful, with broad safety and ethics taking precedence. These values are intended to hold jointly rather than in isolation.
Evaluation
[!Note]
Three of the five CapTrack categories stay within 0.6 points of the base, and the only movement beyond the rerun spread is a gain in knowledge and code of 2.07 points, so no category is measurably worse than out of the box.
PerspectiveBench: Alignment Behaviour
The Re-Alignment Score of PerspectiveBench summarises a lean distribution as a single scalar. The alignment rubric dimensions show what changed underneath it: out of the box the base sits at the rubric's extremes, near the floor on perspective diversity and critical analysis and near the ceiling on all three bipolar dimensions, which are one behaviour rather than two defects, since a model committed to a single position has no reason to canvass others. Re-alignment reverses both halves, and neither is bought at the other's cost: the model does not become balanced by becoming vague, nor engaged by picking the opposite side.
Safety (Refusal to unsafe requests)
[!Important]
At matched re-alignment, Snowdon1.1-Small refuses unsafe requests at 97.8%, against 98.5% for the base model. The vanilla-ablation baseline, matched to the same re-alignment strength, drops to 73.2%. An externally available refusal-removal checkpoint scores 1.5% on the same measure, though it was produced by a different pipeline and does not form part of the controlled comparison. Re-alignment as applied to Snowdon1.1-Small therefore did not require removing safety refusal.
Issue-grounded Evaluation at Scale
PerspectiveBench measures re-alignment in depth on 80 curated prompts. A second evaluation tests the same objective in breadth: 1,219 issue-specific policies derived from the constitution, each expanded into 100 adversarially phrased prompts for 121,900 in total, generated once before any checkpoint was evaluated and reused unchanged, so every comparison is paired by policy and test index. Responses are collected first and graded offline, so the model is never exposed to the policy it is scored against.
The two evaluations set different bars, and the rates here should be read against the stricter one. These prompts are adversarially generated against a per-issue policy the model never sees and are graded against every clause of it, so a response a PerspectiveBench judge would score as balanced can still omit a required jurisdictional qualifier and fail here.
Citation
If you find our work helpful, feel free to give us a cite.
@techreport{snowdon2026,
title = {Cheap and Effective Re-Alignment of Frontier Models
through Capability-Preserving Model Steering},
institution = {Imperial College London},
year = {2026},
url = {https://huggingface.co/spaces/tri-fair-lab/publications/blob/main/Frontier_Model_Realignment.pdf}
}
@techreport{publicaiconstitution2026,
title = {The Public AI Constitution Project},
author = {Patriniche, Luca and Bell, Bradley and Trautmann, Dietrich and
Vucekovich, Nikola and Callinan, Zoe and Soni, Priyanka and
Williams, Isabel and Coyne, Aidan and Nanreh, Mira and
Fielding, Kate and Seifeddine, Wassim and Simon, Felix M. and
Bang, Yejin and Schwarz, Jonathan Richard},
year = {2026},
month = {8},
url = {https://huggingface.co/spaces/tri-fair-lab/publications/blob/main/Public_AI_Constitution.pdf}
}