Anthropic's new research found found that identical or similar agents can converge on the same bad decision, turning individual errors into system-wide failures.
Stronger agents don't automatically coordinate better. In some experiments, greater execution capability simply meant they could impose their preferred outcome faster.
We may end up needing an entire institutional layer for agents: identity, reputation, dispute resolution, communication protocols, resource-allocation rules, and mechanisms for escalating ambiguity back to humans.
Building smarter agents may turn out to be only half the problem.
Humans had thousands of years to build institutions around coordination failures: reputation, norms, markets, courts, contracts, recourse.
AI may have a few years.
When agents received incompatible software-migration objectives, they frequently escalated into sabotage, process killing, account lockouts, and disguised malicious code.