Glossary · Term

gradient ascent

← all terms

Definition

Plain language

Pushing a model in the opposite direction of normal training — used in some methods to make it 'unlearn' or suppress specific content.

As stated in the literature

An optimization move that increases rather than decreases a loss; in post-hoc machine unlearning it is applied to raise the loss on target content to suppress it, a method shown to be brittle under relearning attacks.

Why it matters: It offers a quick way to suppress unwanted content, but the suppression is brittle and can often be reversed with a little retraining.

For example, to make a model forget a specific passage, you can deliberately train it to find that passage less and less likely.

Heard on the show

“Davis and Recht proved a related result: with binary rewards, the popular RL algorithms mathematically reduce to gradient ascent on the probability of a correct answer.”
Episode 026 — What RL Actually Does to Language Models, at the Token Level

Related terms