Concept · 8 episode(s)

Post-Training

← all concepts

Definition

Post-training is everything you do to a model after pretraining: SFT, RLHF, DPO, preference tuning, safety training, tool-use training. Most of what users actually experience as “model behavior” comes from post-training, not pretraining.

Episodes covering this

  1. 261
    Split the Same Story Across Five Messages and the Model Switches Sides
    Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
    Wu, Wang, Chen et al. · The Hong Kong University of Science and Technology (Guangzhou)·23 min·Sep 05, 2026
  2. 246
    160 Perfect Refusals, And The Refusals Were The Leak
    Inadvertent Context Leakage in Language Models
    Fairoze, Mangaokar, Chaudhuri et al. · University of California·20 min·Aug 21, 2026
  3. 219
    Forty-Four AI Models, One Word, And The Newest Ones Conform Most
    The One-Word Census: Answer-Choice Conformity Across 44 Language Models
    Parikh · Cornell Tech·15 min·Jul 15, 2026
  4. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
    Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
    Wei, Kim · Princeton University·22 min·Jun 23, 2026
  5. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
    Natively Unlearnable Large Language Models
    Ghosal, Maini, Raghunathan · Carnegie Mellon University·22 min·Jun 15, 2026
  6. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
    Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
    Li, Zhan, Zhang et al. · Shanghai AI Laboratory / The Chinese University of Hong Kong·31 min·May 16, 2026
  7. 026
    What RL Actually Does to Language Models, at the Token Level
    Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
    Akgül, Kannan, Neiswanger et al. · University of Southern California·24 min·May 08, 2026
  8. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
    Emotion Concepts and their Function in a Large Language Model
    Sofroniew, Kauvar, Saunders et al. · Anthropic·22 min·May 02, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.