Abliteration

FrontierGovernance, Safety and Ethics

Abliteration is a technique for removing an open-weight AI model's built-in refusal behaviour by directly altering the internal pathway responsible for it, rather than retraining the whole model. The name is a blend of "ablation", surgically removing a part, and "obliteration", destroying it entirely. Because open-weight models can be downloaded and modified by anyone, this makes safety guardrails removable after release, not just before it.

In practice

A vendor's safety claims about their base model may not hold once that model's open-weight version has been abliterated and redistributed by someone else. If a supplier builds on open weights, their safety case covers the copy they control and nothing else. Ask what they assume about the versions circulating outside it.

Not sure where your organisation stands?

Take the free AI-readiness diagnostic.

Start the diagnostic