Definition
Activation steering is an inference-time technique that adds a fixed vector to a model’s internal activations to push behavior in a desired direction — more honest, more refusing, less sycophantic — without retraining. The steering vector is typically derived by contrasting activations on prompts that do and don’t exhibit the target behavior.
Episodes covering this
Worth reading next
Papers we haven't done a deep dive on yet, but would recommend on this topic.
- Representation Engineering: A Top-Down Approach to AI Transparency
- Steering Language Models With Activation Engineering
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
- Refusal in Language Models Is Mediated by a Single Direction
- Discovering Latent Knowledge in Language Models Without Supervision