Definition
Plain language
Whether a model is able to do something at all, given the right prompting or setup.
As stated in the literature
The maximum performance a model can reach on a task under favorable conditions, contrasted with propensity to do it spontaneously.
Why it matters: Distinguishing what a model can do from what it tends to do is essential for both safety evaluation and product design.
For example, a model might be capable of solving a hard logic puzzle with the right prompt but never spontaneously attempt that level of reasoning by default.
Heard on the show
“Which predicts something you can test on capability, and they do.”Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak
Mentioned in 111 episodes
- 246
- 245
- 244
- 240
- 237
- 232
- 224
- 223
- 221
- 207
- 205
- 202
- 198
- 197
- 195
- 194
- 193
- 192
- 190
- 189
- 187
- 185
- 184
- 183
- 180
- 175
- 172
- 169
- 166
- 163
- 160
- 158
- 157
- 156
- 152
- 150
- 148
- 147
- 146
- 145
- 144
- 143
- 142
- 133
- 132
- 131
- 129
- 128
- 125
- 123
- 118
- 117
- 114
- 112
- 111
- 110
- 108
- 107
- 105
- 104
- 103
- 099
- 094
- 093
- 092
- 090
- 087
- 084
- 083
- 082
- 081
- 077
- 076
- 072
- 071
- 069
- 068
- 067
- 066
- 064
- 063
- 061
- 058
- 057
- 054
- 053
- 052
- 049
- 048
- 047
- 045
- 044
- 040
- 039
- 035
- 034
- 029
- 028
- 026
- 024
- 023
- 021
- 017
- 013
- 011
- 008
- 007
- 004
- 003
- 002
- 001
Related concepts
Agent Scaffolding
Agentic Vuln Discovery
AI Efficiency & Cost
Benchmark Contamination
Capability Elicitation
Capability vs. Efficiency
Capability vs. Propensity
Creation-Audit Loop
Exploit Generation
GDP-Weighted Evaluation
Inference-Time Scaffolding
Knowledge Distillation
Math Benchmarks
Needle-in-a-Haystack
Recursive Agent Optimization
Sandbagging
Seed-and-Amplify
Self-Efficacy
Step Amplification Factor
Structural Transfer
Test-Time Compute
Training Methods
Web Agents
Workflow Search