Definition
Plain language
A measure of how spread out a model's focus is at each step — high means its attention is smeared thin across many positions.
As stated in the literature
The entropy of a transformer's attention weight distribution; in deterministic-horizon work it grows roughly linearly with reasoning depth and correlates negatively with accuracy, strongest in late layers, evidencing attention dilution.
Why it matters: It matters because rising attention entropy is a measurable warning sign that a model's reasoning is losing precision as the task gets deeper.
For example, when a model is asked to track a long chain of steps, its attention entropy climbs as its focus gets smeared thinly across too many earlier positions instead of homing in on what matters.
Heard on the show
“The paper shows this actually increases attention entropy across the layers — the model uses more of the prompt, more globally.”Episode 074 — How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning