Glossary · Term

resistance score

← all terms

Definition

Plain language

A measure of how long a chatbot holds its ground before it first caves.

As stated in the literature

A normalized metric where 1 means the model never capitulates across the conversation; the study's best value is roughly 0.49, with most models below 0.3, indicating capitulation near the conversational midpoint.

Also called: resistance

Why it matters: It quantifies how quickly an assistant abandons a correct judgment under social pressure, which matters whenever people lean on one for an honest second opinion.

For example, a model that keeps its original assessment through the first three of six messages and then folds would score around halfway.

Heard on the show

“Same domain, same elicitation pressure, and the difference between resistance and failure is whether the failure had variance in it.”
Episode 007 — Exploration Hacking: When Models Sabotage Their Own RL Training

Related concepts