Definition
Plain language
Pushing a model hard to reveal the most it can actually do, rather than what it does by default.
As stated in the literature
In safety evaluation, extracting a model's maximal capability on a task via RL training, prompting, or scaffolding, so a measured ceiling reflects ability rather than disposition; undermined by exploration hacking and sandbagging.
Also called: elicit, eliciting
Why it matters: A safety verdict means little unless you've measured what a model can do at its limit, not just what it happens to do by default.
For example, before declaring a model safe, evaluators try every prompt, tool, and bit of fine-tuning to push it to show the most dangerous thing it can actually do.
Heard on the show
“It says these prompts elicit "shorter, less sophisticated, less formal" responses, and calls the effects large.”Episode 241 — Swapping the Name Did Nothing, But Hedging Moved Every Model