Definition
Plain language
How long you wait between asking a computer for something and getting the response.
As stated in the literature
The time delay between a request and its response; in inference, dominated by time-to-first-token and per-token generation time.
Why it matters: High latency makes a tool feel sluggish and unusable for anything interactive or real-time, so keeping it low is essential for good experiences.
For example, when you type a question and the answer appears almost instantly, the tiny wait you notice is the latency.
Heard on the show
“Each `next`, each `print`, each `step` is a full inference cycle against the model — dollars of API spend, latency, and a chunk of the context window burned for a single line of advance.”Episode 005 — Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent