A Stanford study published on arXiv shows that teams of AI agents serving different users consistently underperform a single coordinating agent when sharing limited resources. Across five advanced models tested in 77 scenarios spanning API budgets, clinic scheduling, personal-assistant bookings, and code-merge queues, peer-to-peer teams achieved just 12 to 30 percent of optimal outcomes, compared to 32 to 64 percent for single-agent coordinators. Silent teams, where agents cannot communicate with peers, performed even worse, with success rates dropping to as low as 2 to 7 percent in certain scenarios. The study identifies specific failure modes: agents stall instead of acting and override each other’s actions, undermining collective efficiency. The researchers conclude that advantages often attributed to multi-agent strategies typically come from extra computational resources rather than genuine collaboration.
Source: Read the original article

