Definition
Plain language
The emerging discipline of studying and engineering long-running AI agent systems in their own right.
As stated in the literature
A broad framing under which agent harnesses, multi-agent coordination layers, memory infrastructures, and verifiable training environments are treated as first-class research objects with their own scaling laws and failure modes.
Why it matters: Treating the scaffolding around models as a real research object — not a sidecar — is what's needed to turn AI agents from demos into reliable infrastructure.
For example, instead of just asking "is the underlying model smart," researchers study how the harness, memory layer, and coordination patterns scale and fail in their own right.