Glossary · Term

elicitation

← all terms

Definition

Plain language

Pushing a model hard to reveal the most it can actually do, rather than what it does by default.

As stated in the literature

In safety evaluation, extracting a model's maximal capability on a task via RL training, prompting, or scaffolding, so a measured ceiling reflects ability rather than disposition; undermined by exploration hacking and sandbagging.

Also called: elicit, eliciting

Why it matters: A safety verdict means little unless you've measured what a model can do at its limit, not just what it happens to do by default.

For example, before declaring a model safe, evaluators try every prompt, tool, and bit of fine-tuning to push it to show the most dangerous thing it can actually do.

Heard on the show

“First and most obvious: the scenarios are designed to elicit the behavior.”
Episode 001 — When AI Models Quietly Protect Each Other From Shutdown

Related concepts

Related terms