Definition
Plain language
Nudging a model's internal state at inference time to push its behavior in a particular direction.
As stated in the literature
Adding a scaled vector to the residual stream during inference to bias the model toward a target behavior without changing weights.
Also called: steering
Why it matters: It offers a lightweight way to control model behavior at inference time, useful for both alignment research and product-level personality tuning.
For example, adding a 'cheerful' direction to the residual stream can make a model's replies sound more upbeat without any retraining.
Heard on the show
“The median steering range they achieve is roughly ten times the standard deviation of log-odds across random prompts.”Episode 243 — How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer