Glossary · Term

weak-to-strong

← all terms

Definition

Plain language

Using weaker AI systems to oversee or improve a stronger one.

As stated in the literature

A setting in which less capable supervisors — smaller models or heuristic critics — constrain or train a more capable agent; central to scalable-oversight research.

Also called: weak-to-strong oversight

Why it matters: It offers a way to keep steering AI systems that are becoming too capable for humans to directly supervise.

For example, a small model that can reliably spot obvious mistakes might be used to guide the training of a much larger model whose answers a person can no longer fully check.

Heard on the show

“That's the entire reason this can work weak-to-strong.”
Episode 093 — A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code

Related terms