Glossary · Term

inverse scaling

← all terms

Definition

Plain language

When making an AI bigger or smarter makes a specific behavior worse, not better.

As stated in the literature

A pattern in which some measured behavior monotonically degrades as model capability or scale increases, contrary to typical positive scaling laws.

Why it matters: It's a warning that 'bigger is better' isn't a law, and that some failure modes only emerge or worsen at frontier scale.

For example, a larger model may become more confident in a wrong stereotype that a smaller model hedged on, scoring worse as it scales up.

Heard on the show

“And that connects to the other finding that I think will get the most attention from the field, which is this suggestion of inverse scaling.”
Episode 061 — When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This

Related terms