Definition
Plain language
A region of a model's possible internal states that pulls nearby states toward a common settled point.
As stated in the literature
In iterative or recurrent inference, a region of state space whose dynamics converge to a single fixed point, used to describe regimes where a token's hidden state stays stable across iterations.
Also called: basin, attractor basins
Why it matters: Understanding attractor basins explains why iterative models converge to stable answers — and why a small nudge can sometimes flip them dramatically.
For example, once a token's hidden state lands in a 'this is a question' basin, repeated refinement keeps pulling it back to the same configuration.
Heard on the show
“The matched-baseline gap is real, the basin-shift analysis is mechanistically illuminating, and the position-zero finding is genuinely surprising regardless of how it generalizes.”Episode 032 — A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking