Glossary · Term

long-context

← all terms

Definition

Plain language

Models or tasks where the input can be tens or hundreds of thousands of words long.

As stated in the literature

The regime where transformer input sequences are large enough that attention cost, memory pressure, and degradation of recall become primary engineering and modeling concerns.

Why it matters: It's the setting where most engineering tradeoffs in modern LLM serving and architecture design currently play out.

For example, asking a model to answer questions about a 200-page contract in a single prompt puts it firmly in the long-context regime.

Heard on the show

“Google made noise about long-context Gemini subsuming retrieval two years ago.”
Episode 198 — The Model That Knows the Answer and Can't Say It

Mentioned in 11 episodes

  1. 198
    The Model That Knows the Answer and Can't Say It
  2. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  3. 130
    Why AI Agents Coordinate Better Through a Shared Board Than a Boss
  4. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  5. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  6. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  7. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  8. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  9. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  10. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  11. 016
    Why Your Coding Agent Stalls While the GPU Runs Hot

Related concepts

Related terms