Guide · 43 episodes · updated 2026-09-06

Self-correction: when models actually fix their own mistakes, and when they just say so

← all guides

Can AI models reliably catch and correct their own errors without external feedback?

Self-correction asks a model to critique and revise its own output, ideally without anyone else pointing out what went wrong. The papers keep circling back to it because the promise is obvious and the reality is patchy. Several episodes show real recovery: that -cause a crash and retry, robots that abandon a failed grasp, systems that rewrite their own buggy code or stale notes. Others complicate the picture, arguing that revision without new information plateaus, that a 'reflection condition' meant to curb bad behavior mostly fails, and that phrases like 'wait, let me check' can appear after the answer is already locked in, making the correction look real while doing nothing. The through-line is that spotting a mistake is easier than fixing it.

What self-correction means

Self-correction has a model critique and revise its own output, ideally fixing errors without external feedback. The empirical story is mixed: models are decent at spotting their own mistakes when prompted, less reliable at correcting them, and prone to second-guessing correct answers.

The episodes (43)

Newest first. Each line is what that paper contributed to the question.

Papers we have not covered yet

Other guides

Intro written by Anthropic's Claude Sonnet 5; episodes selected and edited by Garrett Casey. Episode notes come from each episode's own analysis. How PaperDive is made.