Glossary · Term

cache read

← all terms

Definition

Plain language

Reusing text a model has already processed, which the provider charges less for than fresh text.

As stated in the literature

Prompt-cache hits billed at a discounted per-token rate; can dominate token accounting in long agentic runs with large stable prefixes.

Also called: cache reads

Why it matters: In long agent runs most of the tokens are repeated context, so whether they count as cache reads can change the bill by an order of magnitude and distort any cost comparison that ignores them.

For example, an agent that re-sends the same 50-page project brief with every step pays full price the first time and a discounted rate for that same text on each later call.

Heard on the show

“Over ninety-seven percent of those were cache reads, which is reused context charged at a discounted rate.”
Episode 285 — What a Perfect Score Hides: Auditing an AI Agent That Scored 100

Related terms