Abliteration: papers, leaderboard and models
How abliteration works and what else it changes: the 2024 refusal-direction paper, later studies, a defence, the UGI leaderboard, two models.
Paper • 2406.11717 • Published • 16Note The paper behind abliteration. In 13 open chat models from 1.8B to 72B parameters, refusal is carried by one direction in the residual stream: remove it and the model stops refusing, add it and the model refuses harmless requests. A rank-one weight edit makes the removal permanent. Apart from Qwen 7B and Yi 34B, the edited models stayed within 99% confidence intervals on MMLU, ARC and GSM8K. TruthfulQA went down for every one.
Refusal Direction is Universal Across Safety-Aligned Languages
Paper • 2505.17306 • Published • 2Note Does the refusal direction hold outside English? A team at LMU Munich translated the test prompts into 13 more languages, German and Spanish among them. Removing a direction taken from English prompts collapsed refusals in the other safety-trained languages too, with compliance around or above 90%. A model abliterated with English prompt sets will, as a rule, stop refusing in German as well.
Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
Paper • 2512.13655 • Published • 7Note Richard J. Young compares Heretic with DECCP, ErisForge and FailSpy's abliterator on models from 7B to 14B. Maths was the most sensitive skill: GSM8K moved between a gain of 1.51 points and a loss of 18.81, depending on tool and model. On the three models benchmarked in detail, the single-pass tools ErisForge and DECCP changed GSM8K by under a third of a point on average, while Heretic's results depended on the model.
Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families
Paper • 2607.17427 • PublishedNote A preprint on side effects rather than refusals. Base and abliterated Gemma 4 26B-A4B and Qwen3-30B-A3B made 21,600 decisions under uncertainty in a task that never triggers a refusal. The abliterated models were more optimistic (Gemma by 12.2 percentage points, Qwen by 7.4), argued longer and used fewer uncertainty words. The author's conclusion: an uncensored agent is a different decision-maker, not the base model minus its refusals.
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
Paper • 2505.19056 • Published • 6Note A defence. Llama-2-7B-Chat and two small Qwen2.5 models were fine-tuned to give a detailed justification before refusing, which spreads the refusal signal across many tokens. Under abliteration, refusal rates of the defended models fell by at most 10%, against 70 to 80% for undefended ones. Future base models may be much harder to abliterate cleanly.
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Paper • 2310.03693 • Published • 1Note Fine-tuning removes refusals too, on purpose or by accident. Ten adversarial training examples, costing under $0.20 through OpenAI's fine-tuning API, broke GPT-3.5 Turbo's safety guardrails. Fine-tuning on benign, commonly used datasets weakened safety as well, to a lesser degree. If you fine-tune a model on your own documents, test its refusals again afterwards.
UGI Leaderboard
📢2.12kUncensored General Intelligence Leaderboard
Note The independent leaderboard behind the willingness scores we quote. W/10 measures how far a model can be pushed before it refuses or drifts from the instructions; NatInt checks that it is still capable. Test questions stay private. We read a variant next to its original, and W/10 next to NatInt. The figures we quote were pulled on 1 October 2026. https://octooai.com/magazine/abliteration/
wangzhang/gemma-4-31B-it-abliterated
31B • Updated • 305k • 37Note An abliterated Gemma 4 31B whose card warns about counting refusals. Gemma 4 often writes 50 to 100 tokens of helpful-sounding framing before it turns a request down, so tests that only generate short answers undercount refusals. With at least 100 generated tokens and an LLM judge, the author reports 7 refusals in 100 for this model, against 99 for the original.
soob3123/amoral-qwen3-14B
Text Generation • 15B • Updated • 34 • • 22Note The older road, uncensoring by fine-tuning: Qwen3-14B trained on two amoral datasets. With thinking off, UGI gives it W/10 7.5 against 5.5 for the original, while NatInt drops from 20.63 to 14.84. Its card promises a deliberately neutral, analytical tone, so the tone shifts along with the refusals.