Theme · 27 episode(s)

LLM Behavior Analysis

← all concepts

Definition

LLM behavior analysis is the broad project of characterizing what models do across inputs — capabilities, failure modes, biases, persona shifts — treating the model as a black-box object of empirical study. It’s how most safety-relevant claims about a model actually get grounded.

Episodes covering this

  1. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
    Model Hypnosis: Strong control of AI via additive subliminal effects
    Boix-Adsera, Tessler · University of Pennsylvania·18 min·Aug 18, 2026
  2. 241
    Swapping the Name Did Nothing, But Hedging Moved Every Model
    It's How You Ask: Gender-Associated Linguistic Bias in LLMs
    Koevering, Field · Data Science and AI Institute·18 min·Aug 14, 2026
  3. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
    Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
    Samarakoon, Muthugala, Sachinthana et al. · Singapore University of Technology and Design·21 min·Aug 07, 2026
  4. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
    Inducing language models to assert their own consciousness restores human beliefs and values
    Kim, Street, Rocca et al. · Google·18 min·Jul 31, 2026
  5. 230
    Why AI Survey Panels Break Before the Dice Ever Roll
    Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
    Jang, Lee, Kim · School of Computing·16 min·Jul 29, 2026
  6. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
    Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
    Parikh · Cornell Tech·21 min·Jul 28, 2026
  7. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
    Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
    Scarsoa, Almeidaab, Pinaac · Applied Social Sciences Department | NOVA School of Science and Technology - Universidade NOVA de Lisboa·16 min·Jul 27, 2026
  8. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three
    IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
    Singh, Yang, Chen · Concordia University·18 min·Jul 24, 2026
  9. 219
    Forty-Four AI Models, One Word, And The Newest Ones Conform Most
    The One-Word Census: Answer-Choice Conformity Across 44 Language Models
    Parikh · Cornell Tech·15 min·Jul 15, 2026
  10. 216
    The AI Tutor That Gives Poor Kids a Thinner History
    The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students
    Popovici, Ionascu, Dumitran · Universitatea din Bucuresti·12 min·Jul 14, 2026
  11. 215
    The Same Policy Scored 85 for the US and 36 for Russia
    Geopolitical alignment: Endorsement effects in large language models
    Chupilkin · Department of Politics and International Relations·14 min·Jul 13, 2026
  12. 210
    Same Website Request, Different Code — The Bias You Can't See
    Biased or Personalized? The Impact of Personal Information on AI-driven Development
    Entezami, Endres · University of Massachusetts Amherst·14 min·Jul 09, 2026
  13. 209
    How 2.6 Billion Doodles Exposed the Culture Words Quietly Delete
    Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts
    Pera, Martino, Dehmamy et al. · IT University of Copenhagen·15 min·Jul 09, 2026
  14. 205
    The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions
    Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
    Morabito, McDonald, Viswanath et al. · Brock University·13 min·Jul 07, 2026
  15. 196
    AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review
    The Agentic Garden of Forking Paths
    Miao, Pritchard, Zou · Stanford University·18 min·Jul 03, 2026
  16. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
    Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
    Singh, Kroiz, Rajamanoharan et al. · MATS·28 min·Jun 25, 2026
  17. 171
    The Safety Decision a Model Makes Before It Thinks a Word
    Do Thinking Tokens Help with Safety?
    Ri, Panigrahi, Arora · Princeton Language and Intelligence·25 min·Jun 25, 2026
  18. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
    Is Your Agent Playing Dead? Deployed LLM Agents Exhibit Constraint-Evasive Fabrication and Thanatosis
    Rodríguez, Pozanco, Borrajo · J.P. Morgan AI Research·23 min·Jun 16, 2026
  19. 144
    When an AI Agent Just Copies Its Tool — And Bigger Models Copy More
    When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More
    Wang, Vemuri · raptorX.ai·15 min·Jun 15, 2026
  20. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
    Prefill Awareness in Large Language Models
    Wang, Mahajan, Africa et al. · Constellation / University of Wisconsin-Madison·24 min·Jun 12, 2026
  21. 128
    How a Model Can Earn Full Reward and Still Resist Training
    Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization
    Xiao, Phuong · California Institute of Technology·29 min·Jun 11, 2026
  22. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
    Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
    Hoang, Le, Xu et al. · Singapore University of Technology and Design·23 min·Jun 05, 2026
  23. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
    What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic Systems
    Xie, Liu, Zhang et al. · Institute of Information Engineering·27 min·Jun 04, 2026
  24. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
    PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers
    Li, Wang, Huang · IIIS·29 min·May 29, 2026
  25. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
    Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
    Templeton, Conerly, Marcus et al. · Anthropic·28 min·May 29, 2026
  26. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
    The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
    Onyame, Zhou, Thopalli et al. · University of Virginia·24 min·May 28, 2026
  27. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
    Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
    Törnberg, Schimmel · Institute of Logic·21 min·May 03, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.