Glossary · Term

activation steering

← all terms

Definition

Plain language

Nudging a model's internal state at inference time to push its behavior in a particular direction.

As stated in the literature

Adding a scaled vector to the residual stream during inference to bias the model toward a target behavior without changing weights.

Also called: steering

Why it matters: It offers a lightweight way to control model behavior at inference time, useful for both alignment research and product-level personality tuning.

For example, adding a 'cheerful' direction to the residual stream can make a model's replies sound more upbeat without any retraining.

Heard on the show

“The median steering range they achieve is roughly ten times the standard deviation of log-odds across random prompts.”
Episode 243 — How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer

Mentioned in 32 episodes

  1. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  2. 241
    Swapping the Name Did Nothing, But Hedging Moved Every Model
  3. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  4. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  5. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
  6. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
  7. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
  8. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
  9. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  10. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  11. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  12. 205
    The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions
  13. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
  14. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  15. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  16. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  17. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  18. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  19. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  20. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  21. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  22. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
  23. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  24. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  25. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  26. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  27. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  28. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  29. 026
    What RL Actually Does to Language Models, at the Token Level
  30. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  31. 018
    Language Models Compute the Rational Move, Then Override It
  32. 006
    What Happens Inside Claude When It Decides to Blackmail Someone

Related concepts

Related terms