Glossary · Term

post-training

← all terms

Definition

Plain language

The extra rounds of training that turn a finished base model into something useful and well-behaved, after the big initial training on raw text.

As stated in the literature

The stages applied after pretraining — supervised fine-tuning, RLHF, instruction tuning, and related methods — that adapt a base language model into an assistant or specialist; where alignment and most behavioral shaping happen.

Also called: post-trained, post-train

Why it matters: It is where most of a model's helpfulness, manners, and safety come from, so it largely determines how the finished assistant behaves.

For example, post-training is the stage that turns a raw text-predictor into a chatbot that politely answers questions and refuses harmful requests.

Heard on the show

“They take an open model, OLMo-3-32B-Think, and walk it through post-training stage by stage.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 20 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 242
    Making a Vision Model Better by Showing It Blurry Images
  3. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  4. 219
    Forty-Four AI Models, One Word, And The Newest Ones Conform Most
  5. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
  6. 193
    Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer
  7. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  8. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  9. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  10. 092
    When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks
  11. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  12. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  13. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  14. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  15. 026
    What RL Actually Does to Language Models, at the Token Level
  16. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  17. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  18. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  19. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  20. 004
    The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms