Definition
Plain language
A quick edit to a downloaded AI model that deletes its ability to say no.
As stated in the literature
Training-free removal of refusal behavior by estimating a single refusal direction in the residual stream from mean activation differences on harmful versus harmless prompts, then projecting that component out of every weight matrix that writes into the stream; it preserves general capability and runs in minutes on consumer hardware.
Also called: abliterated, abliterate, abliterating
Why it matters: It means the safety behavior built into publicly released weights can be stripped out cheaply by almost anyone, so any release decision has to assume the refusals will not survive.
For example, someone downloads an open model that normally declines to explain how to make a weapon, runs a short script on their gaming PC, and gets back a version that answers every request without hesitation.
Heard on the show
“So, the cheapest version of the attack is called abliteration, and its specific character shapes everything that follows.”Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers