Glossary · Term

Universal Transformers

← all terms

Definition

Plain language

A Transformer variant that runs the same layers over and over to think harder, rather than stacking more layers.

As stated in the literature

A 2018 architecture that applies a single Transformer block recurrently across depth, sharing weights across iterations; foundational reference for depth-recurrent and looped sequence models.

Why it matters: It introduced the idea that adaptive depth — thinking more on harder inputs — can substitute for stacking more parameters.

For example, a Universal Transformer might run its single layer five times on an easy input and twenty times on a hard one to think longer.

Heard on the show

“Universal Transformers, recurrent-depth models.”
Episode 032 — A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking

Related terms