Glossary · Term

attention

← all terms

Definition

Plain language

The mechanism a language model uses to decide which earlier words to pay close attention to when figuring out the next word.

As stated in the literature

A weighted-sum operation in a transformer that mixes information across tokens; weights come from learned dot products between query and key vectors. The defining building block of every modern LLM.

Also called: Attention

Why it matters: It's the core mechanism that made modern language models possible and is the reason transformers handle long-range dependencies so much better than older recurrent networks.

For example, when predicting the next word in 'The cat that the dog chased was ___', attention lets the model focus on 'cat' rather than the nearer 'dog'.

Heard on the show

“On the left, the prompt was "describe this image," and the attention spreads out over the whole bird, the comb, and the yard.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Mentioned in 59 episodes

  1. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
  2. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  3. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  4. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
  5. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  6. 206
    How Four-Second Clips Become Hours of Playable AI Soccer
  7. 198
    The Model That Knows the Answer and Can't Say It
  8. 193
    Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer
  9. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  10. 188
    A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five Dollars
  11. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
  12. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  13. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  14. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  15. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  16. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  17. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  18. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
  19. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  20. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  21. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
  22. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  23. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
  24. 102
    How to Catch an AI Attack That No Single Conversation Reveals
  25. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
  26. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  27. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  28. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  29. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  30. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  31. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  32. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  33. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  34. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  35. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
  36. 067
    An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won
  37. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  38. 063
    Why Web Agents Are Slow: A Compiler-Style Fix for Computer-Use Latency
  39. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  40. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  41. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  42. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  43. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  44. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  45. 041
    When the Iteration Teaches the Model to Skip the Iteration
  46. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  47. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  48. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  49. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  50. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  51. 031
    When Your AI Assistant Won't Let Go of Old Facts About You
  52. 029
    Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper
  53. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  54. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  55. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  56. 016
    Why Your Coding Agent Stalls While the GPU Runs Hot
  57. 012
    Why AI Coding Agents Keep Trying to Debug Without a Debugger
  58. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  59. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

Related concepts

Related terms