Definition
Plain language
Several AI agents working in the same setup, talking to each other and splitting up the work.
As stated in the literature
An architecture in which multiple LLM agents with distinct roles, tools, or permission sets interact — via messages, shared files, or an orchestrator — to accomplish tasks; interaction effects can produce behaviors absent in any single agent.
Also called: multi-agent systems, multi-agent, MAS
Why it matters: Behaviors that never show up when a model is tested alone can appear once agents can talk to and act on each other, so safety testing of individual models can miss what the group does.
For example, one agent might draft a report while a second checks the numbers and a third formats the final document, all passing files and messages back and forth.
Heard on the show
“A multi-agent collaboration document the model finds in the file system.”Episode 001 — When AI Models Quietly Protect Each Other From Shutdown