Glossary · Term

KV cache

← all terms

Definition

Plain language

The model's short-term memory of everything it has read so far in the current conversation.

As stated in the literature

Stored attention keys and values from previous tokens that a transformer reuses during generation; its size grows linearly with context length and often dominates inference memory.

Also called: KV-cache, key-value cache, KV caches

Why it matters: Managing the KV cache is the single biggest determinant of how much long-context inference costs in memory and dollars.

For example, when generating the 1000th token of a response, the model reuses the cached keys and values from the previous 999 tokens instead of recomputing them.

Heard on the show

“Stack all of those up and you get the KV cache.”
Episode 226 — How a Speed Feature Lets a Stranger Poison Your AI's Answer

Mentioned in 7 episodes

  1. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
  2. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  3. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  4. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  5. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  6. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  7. 016
    Why Your Coding Agent Stalls While the GPU Runs Hot

Related concepts

Related terms