AI glossary
Abliteration
Abliteration is a technique for removing a language model’s ability to refuse requests by finding the single internal direction that represents refusal and erasing it from the model, either while it runs or permanently in its weights. It needs no retraining and no new training data. The result is usually published as an “abliterated” version of an existing open-weights chat model, which answers requests the original model would have declined.
The word is a blend of “ablate” and “obliterate”. It was coined in 2024 by an independent developer who publishes under the name FailSpy; his Hugging Face model card explains it as “Ablate + obliterated = Abliterated”, a play on the research term “ablation” chosen to set these models apart from “uncensored” fine-tunes.
The research behind it
Abliteration is a practical application of an interpretability finding. In April 2024, Andy Arditi, Oscar Obeso, Aaquib Syed, Wes Gurnee and Neel Nanda posted “Refusal in LLMs is mediated by a single direction” on the Alignment Forum. The full paper, with Daniel Paleka and Nina Panickssery added as co-authors, was submitted to arXiv on 17 June 2024 (arXiv:2406.11717) and accepted at NeurIPS 2024.
Its central claim: across 13 popular open-source chat models up to 72B parameters, refusal is mediated by a one-dimensional subspace in the model’s residual stream, the running internal representation a transformer updates layer by layer. Removing that direction stops the model refusing harmful instructions; adding it makes the model refuse harmless ones. The authors describe the resulting weight edit as “a novel white-box jailbreak method” and say the finding underscores “the brittleness of current safety fine-tuning methods”.
FailSpy’s first abliterated models, built on Meta’s Llama 3, cited the Alignment Forum post. The technique reached a wide audience through “Uncensor any LLM with abliteration”, a Hugging Face blog post by ML engineer Maxime Labonne published on 13 June 2024, which credits FailSpy’s notebook and library.
How abliteration works, at a high level
The process has three conceptual stages:
- Find the direction. The model is run on two sets of prompts, one it would refuse and one it would answer. The average difference between its internal activations on the two sets points toward a candidate “refusal direction”.
- Pick the best candidate. Candidates from different layers are compared to find the one that most reliably separates refusing from answering.
- Remove it. The direction is either subtracted from the model’s activations at inference time, or the weights are edited (“orthogonalized”) so the model can no longer write to that direction at all. The second option produces a new set of weights that can be shared like any other model.
Nothing about the architecture changes, and no new knowledge is added. The model simply loses the internal signal that triggered its trained refusals. For the wider context, including how such models have been misused, see our guide to uncensored AI models.
Abliteration vs fine-tuning and jailbreaking
| Abliteration | Uncensored fine-tune | Jailbreak prompt | |
|---|---|---|---|
| What changes | One direction removed from activations or weights | Weights retrained on new data | Nothing in the model; only the input |
| Needs training data | Only contrasting prompts to locate the direction | Yes, a fine-tuning dataset | No |
| Works on hosted APIs | No, needs access to the weights | No, needs the weights | Yes |
| Effect is permanent | Yes, if the weights are edited | Yes | No, per conversation |
A jailbreak tricks a model through its input and can be patched by the provider. Fine-tuning changes behavior by training on new examples, which costs compute and data. Abliteration sits between them: a targeted, permanent edit that only works where the weights are available, which is why it applies to open-weights models and not to closed models served through an API.
What it costs
Removing a direction is not free. In Labonne’s tests on Daredevil-8B, a Llama 3 8B derivative, he reported “a performance drop in the ablated version across all benchmarks”. On the Nous benchmark suite the average fell from 55.87 to 55.06, with TruthfulQA dropping from 59.05 to 57.47. He then ran a short preference-tuning pass (DPO) on the abliterated model, which he says “allowed us to recover most of the performance drop”; the resulting NeuralDaredevil-8B-abliterated scored 55.87, matching the source model, although he notes that the healing pass did not improve GSM8K, a math benchmark.
The bigger cost is that abliteration is indiscriminate. The same direction governs a model’s correct refusals (instructions for malware, for example) and its over-refusals (a benign security question). Erasing it removes both.
How widespread it is
A search for “abliterated” on Hugging Face returned 8,338 models on 5 October 2026. That counts models with the word in their name, not models confirmed to have been abliterated, but it shows the technique is a large, visible part of the open-weights ecosystem rather than a niche experiment.
Defenses and later research
Because the original finding showed safety tuning to be fragile, it prompted work on making refusal harder to remove:
- Refusal feature adversarial training (ReFAT), by researchers from Meta FAIR, Toronto and UCL (arXiv, September 2024; ICLR 2025), argues that many attacks share “a universal mechanism” of ablating the refusal feature, and trains models while simulating that ablation to make them more robust.
- The geometry of refusal (Wollschläger et al., ICML 2025) found that refusal can be mediated by multiple independent directions and multi-dimensional “concept cones”, challenging the single-direction picture.
- Extended-refusal fine-tuning (Abu Shairah et al., KAUST, 2025) trains models to justify their refusals at length, spreading the signal across many tokens. In their tests, refusal rates of defended models fell by at most 10% after abliteration, against 70–80% for undefended baselines.
FAQ
What does abliteration mean?
It is a blend of “ablate” and “obliterate”: a technique that removes a language model’s refusal behavior by erasing the single internal direction responsible for it, without retraining the model.
Who invented abliteration?
The underlying finding comes from Arditi, Nanda and co-authors (Alignment Forum, April 2024; NeurIPS 2024). The name and the first public abliterated models came from the developer FailSpy in 2024, and Maxime Labonne’s Hugging Face post of 13 June 2024 popularized the method.
Does abliteration make a model worse?
Slightly, in the documented tests. Labonne measured small drops across benchmarks after abliteration and recovered most of them with an extra DPO fine-tuning pass. The model also loses all refusal-based safety behavior, including refusals that were correct.
Can abliteration be used on ChatGPT or Claude?
No. It requires direct access to a model’s weights or activations, so it applies to open-weights models such as Llama, Mistral or Qwen, not to closed models served through an API.
Is an abliterated model the same as an uncensored model?
Abliteration is one way to make an “uncensored” model. Others include fine-tuning on data without refusals. FailSpy chose the name specifically to distinguish the technique from those fine-tunes.
Related terms
The full story, including the underground services built on such models: Uncensored AI Models: What They Are and How They’re Made.