Definition
Plain language
Making a trained AI forget something specific — a book, a person's data — without retraining it from scratch.
As stated in the literature
The task of removing the influence of designated training data from a model; post-hoc methods suppress content via fine-tuning but are brittle (recoverable in a few relearning steps), motivating architectures where unlearnability is built in and removal is structural rather than cosmetic.
Also called: unlearning, unlearnable
Why it matters: It promises a practical way to honor data-removal requests, but shallow methods can leave the 'forgotten' content easy to recover, so the removal needs to be genuine.
For example, if a model was trained on a book the author wants removed, machine unlearning aims to strip out that book's influence without rebuilding the model from scratch.
Heard on the show
“And the standard way people do this today is what's called post-hoc unlearning — you take the finished model and fine-tune it to suppress the target content.”Episode 145 — Building Forgetting Into a Language Model With One Extra Line of Code