Glossary · Term

threat model

← all terms

Definition

Plain language

A written-down guess about who might attack you, what they want, and what tools they have.

As stated in the literature

The explicit specification of adversary capabilities, access, and goals against which a defense's guarantee is stated; here it grants full weight access and compute while denying ground-truth verification.

Also called: threat models

Why it matters: A defense only means something relative to a stated attacker, so without one, claims like "this is secure" have no content.

For example, one defense assumes the attacker has the full model and plenty of computing power but no way to check whether the answers they get are actually true.

Heard on the show

“The paper is called Vis-Poison, out of a team led by Rujin Liang, and the threat model is about as stripped-down as these things get.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Mentioned in 15 episodes

  1. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
  2. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  3. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  4. 223
    When Grok Graded Its Own Encyclopedia And Marked Itself Down
  5. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
  6. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  7. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  8. 128
    How a Model Can Earn Full Reward and Still Resist Training
  9. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
  10. 103
    AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee
  11. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  12. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
  13. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  14. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  15. 004
    The Sycophancy Circuit That Survives Alignment Training

Related terms