Glossary · Term

transformer

← all terms

Definition

Plain language

The dominant neural network design behind modern language models, built around attention.

As stated in the literature

The architecture introduced in 2017 combining self-attention with feedforward blocks and residual connections, underlying nearly all current frontier LLMs.

Also called: transformers, Transformer

Why it matters: Nearly every advance in modern language and multimodal AI has come on top of this one architecture, so understanding it underlies understanding the field.

For example, GPT, Claude, and Gemini are all transformer-based, processing each token by attending to every other token in their context.

Heard on the show

“Inside a transformer there's a shared bus that researchers call the residual stream.”
Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Mentioned in 32 episodes

  1. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  2. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  3. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
  4. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  5. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  6. 204
    The Length Estimate Hiding Inside a Word-by-Word Model
  7. 198
    The Model That Knows the Answer and Can't Say It
  8. 193
    Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer
  9. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  10. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  11. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  12. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  13. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
  14. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  15. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  16. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  17. 091
    When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning
  18. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  19. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  20. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  21. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  22. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  23. 041
    When the Iteration Teaches the Model to Skip the Iteration
  24. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  25. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  26. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  27. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  28. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  29. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  30. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  31. 016
    Why Your Coding Agent Stalls While the GPU Runs Hot
  32. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

Related concepts

Related terms