Glossary · Term

context window

← all terms

Definition

Plain language

How much text a model can pay attention to at one time.

As stated in the literature

The maximum sequence length a transformer can attend over during a forward pass, bounded by model architecture and memory.

Also called: context windows

Why it matters: It's the hard upper bound on how much information a model can consider in a single pass, shaping which tasks fit naturally and which need workarounds.

For example, with a 200k-token context window, a model can hold roughly a 500-page book in memory at once when answering questions about it.

Heard on the show

“Then the context window starts filling up.”
Episode 002 — An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

Related concepts

Related terms