Concept · 45 episode(s)
Capability vs. Propensity
Read the guide: what the papers say about Capability vs. Propensity →
Definition
Capability vs propensity separates two questions about a model: can it do X if pushed, and does it tend to do X by default. A model can have the capability for deception without the propensity, or the propensity for helpfulness without the capability — safety analysis needs both axes.
Episodes covering this
Worth reading next
Papers we haven't done a deep dive on yet, but would recommend on this topic.
- Alignment faking in large language models
- Language Models (Mostly) Know What They Know
- Frontier Models are Capable of In-Context Scheming
- Chain-of-Verification Reduces Hallucination in Large Language Models
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Sycophancy to Subterfuge: Investigating Reward Tampering in Language Models
- Sabotage Evaluations for Frontier Models