Glossary · Term

attention head

← all terms

Definition

Plain language

One of many small specialists inside a transformer that decides which earlier tokens to focus on.

As stated in the literature

One of multiple parallel sub-units in a transformer attention layer, each computing its own query-key-value projection over earlier tokens.

Also called: attention heads, heads

Why it matters: Different heads end up specializing in different linguistic patterns, and studying them is the entry point to mechanistic interpretability.

For example, one attention head in a transformer might consistently look at the subject of the previous clause while another tracks matching brackets.

Heard on the show

“… out examples — say, a random coin sequence — it writes almost the exact same sequence every time, heads-tails-heads-tails, tidy and alternating, like a person at a party trying to look random. …”
Episode 230 — Why AI Survey Panels Break Before the Dice Ever Roll

Mentioned in 12 episodes

  1. 230
    Why AI Survey Panels Break Before the Dice Ever Roll
  2. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  3. 198
    The Model That Knows the Answer and Can't Say It
  4. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  5. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  6. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
  7. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  8. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  9. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  10. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  11. 018
    Language Models Compute the Rational Move, Then Override It
  12. 004
    The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms