✂️ Abliteration
The refusal direction is computed by comparing the residual streams between target (harmful) and baseline (harmless) samples.
The hidden states of target modules (e.g., o_proj) are orthogonalized to subtract this refusal direction with a given weight factor.
These weight factors follow a normal distribution with a certain spread and peak layer.
Modules can be iteratively orthogonalized in batches, or the refusal direction can be accumulated to save memory.
Finally, I used a hybrid evaluation with a dedicated test set to calculate the acceptance rate. This uses both a dictionary approach and NousResearch/Minos-v1.
The goal is to obtain an acceptance rate >90% and still produce coherent outputs.