Definition
Plain language
Whether a model is able to do something at all, given the right prompting or setup.
As stated in the literature
The maximum performance a model can reach on a task under favorable conditions, contrasted with propensity to do it spontaneously.
Why it matters: Distinguishing what a model can do from what it tends to do is essential for both safety evaluation and product design.
For example, a model might be capable of solving a hard logic puzzle with the right prompt but never spontaneously attempt that level of reasoning by default.
Heard on the show
“Or worse, if the underlying dynamics deepen with capability.”Episode 001 — When AI Models Quietly Protect Each Other From Shutdown
Related concepts
Agent Scaffolding
Agentic Vuln Discovery
AI Efficiency & Cost
Benchmark Contamination
Capability Elicitation
Capability Laundering
Capability vs. Efficiency
Capability vs. Propensity
Creation-Audit Loop
Exploit Generation
GDP-Weighted Evaluation
Inference-Time Scaffolding
Knowledge Distillation
Math Benchmarks
Model Extraction
Needle-in-a-Haystack
Recursive Agent Optimization
Sandbagging
Seed-and-Amplify
Self-Efficacy
Step Amplification Factor
Structural Transfer
Test-Time Compute
Training Methods
Web Agents
Workflow Search