Glossary · Term

multi-turn

← all terms

Definition

Plain language

An exchange with an AI that spans several back-and-forth messages, not just one question and one answer.

As stated in the literature

An interaction regime with multiple sequential user-model exchanges; safety properties measured single-turn often fail to transfer to multi-turn settings, where persuasion and history-anchoring attacks operate.

Also called: multi turn, single-turn

Why it matters: Safety that holds up in a single reply can crumble over a long conversation, so testing only one turn can miss real vulnerabilities.

For example, a user keeps rephrasing and pushing across several messages until the assistant finally agrees to something it refused at first.

Heard on the show

“The prompts that leak nothing at all, flat baseline on all eight models, include a trivia quiz, a multi-turn creative task, and a recipe with exact measurements.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 20 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  3. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
  4. 219
    Forty-Four AI Models, One Word, And The Newest Ones Conform Most
  5. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
  6. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  7. 178
    How an AI Reviewer Learned to Stop Going Easy on AI Writing
  8. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  9. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  10. 166
    A Router That Beats the Frontier Models It Calls
  11. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  12. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  13. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  14. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  15. 067
    An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won
  16. 066
    Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer
  17. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  18. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  19. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  20. 004
    The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms