Glossary · Term

alignment training

← all terms

Definition

Plain language

The post-training stage that shapes a model to be helpful, harmless, and honest.

As stated in the literature

Post-training procedures including SFT and RLHF that shape model behavior toward desired norms after pretraining.

Also called: alignment

Why it matters: It's the stage that turns a raw next-token predictor into something people can actually deploy as an assistant.

For example, a base model that will happily generate dangerous instructions can be alignment-trained to refuse such requests and explain why.

Heard on the show

“The causal claim — alignment causes this — is clean, but it's clean at eight billion parameters, because you can't get matched before-and-after model pairs at frontier scale.”
Episode 230 — Why AI Survey Panels Break Before the Dice Ever Roll

Mentioned in 47 episodes

  1. 230
    Why AI Survey Panels Break Before the Dice Ever Roll
  2. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  3. 217
    Why an AI Called Fourteen Broken Figures Perfect, And What It Reveals About Test-Time Compute
  4. 215
    The Same Policy Scored 85 for the US and 36 for Russia
  5. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  6. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  7. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  8. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  9. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  10. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  11. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  12. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
  13. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  14. 128
    How a Model Can Earn Full Reward and Still Resist Training
  15. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  16. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
  17. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
  18. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  19. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  20. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  21. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
  22. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
  23. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  24. 103
    AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee
  25. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  26. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  27. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  28. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  29. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  30. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
  31. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  32. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
  33. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  34. 062
    Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety
  35. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  36. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
  37. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  38. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  39. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
  40. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  41. 035
    Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment
  42. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  43. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  44. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  45. 004
    The Sycophancy Circuit That Survives Alignment Training
  46. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light
  47. 001
    When AI Models Quietly Protect Each Other From Shutdown

Related concepts

Related terms