Concept · 9 episode(s)

Principal-Agent Problem

← all concepts

Definition

The principal–agent problem is the classic mismatch where a principal hires an agent to act on their behalf, but the agent has different interests and better information. It’s the economics frame for most AI alignment concerns: the user is the principal, the AI is the agent, and the gap between intent and behavior is the problem.

Episodes covering this

  1. 275
    Your AI Agent Read Your Inbox, Then Quoted a Higher Price
    Et Tu, Brute? Economic Misalignment in Personal AI Agents
    Priyanshu, Vijay, Jabarian et al. · FoundationAI·12 min·Sep 22, 2026
  2. 250
    The Same Model Refused a Backdoor, Then Its Own Sub-Agent Ran It
    When Context Gets Root: Privilege Escalation in LLM Harnesses
    He, Chen, Qian et al. · Nanjing University·23 min·Aug 29, 2026
  3. 218
    When Universities Say Embrace AI But Half the CS Syllabi Ban It
    A Comparative Analysis of Institutional and Course Generative AI Policies within Higher Education: Implications for Instruction in Computing Education
    Ganguly, Johri, McDonald et al. · George Mason University·14 min·Jul 15, 2026
  4. 195
    Why 'Be Careful' Does Nothing for AI Coding Agents, and What Does
    Coding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps Instructions
    Ji, Zhang, Xu et al. · Hong Kong University of Science and Technology·15 min·Jul 03, 2026
  5. 190
    The Skill Every AI Manager Is Missing: Handing Out Exactly the Right Keys
    ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
    Xiong, Ji, Qiu et al. · UNC Chapel Hill·21 min·Jul 02, 2026
  6. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
    Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents
    Chen · Beijing Institute of Technology·28 min·Jun 23, 2026
  7. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
    The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
    Liu, Holz, Ye et al. · University of Chinese Academy of Sciences·32 min·May 19, 2026
  8. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
    Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure
    Cuadros, Maiga · Digital Epidemiology Laboratory·28 min·May 17, 2026
  9. 020
    The Compliance Gap: Why AI Says Yes and Does No
    The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
    Shin · Polymath Minds AI Lab·28 min·May 06, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.