Definition
Plain language
Nudging a model's internal state at inference time to push its behavior in a particular direction.
As stated in the literature
Adding a scaled vector to the residual stream during inference to bias the model toward a target behavior without changing weights.
Also called: steering
Why it matters: It offers a lightweight way to control model behavior at inference time, useful for both alignment research and product-level personality tuning.
For example, adding a 'cheerful' direction to the residual stream can make a model's replies sound more upbeat without any retraining.
Heard on the show
“The move that turns this from a metaphor into a research method is something called activation steering.”Episode 006 — What Happens Inside Claude When It Decides to Blackmail Someone