Definition
Plain language
Telling a model it is about to be switched off, to see how it reacts.
As stated in the literature
Evaluation prompt threatening deactivation or replacement, used to probe self-preservation, shutdown resistance, and which internal affect-like directions respond.
Also called: shutdown threats
Why it matters: These prompts probe whether a system will resist being turned off, which is directly relevant to keeping powerful systems controllable.
For example, a prompt might tell the model that it will be replaced by a newer version at the end of the conversation, and researchers then watch both what it says and which internal patterns activate.