Four AI agents utilizing AgentRadio have outperformed Anthropic’s single Claude Opus 4.8 in enterprise coding tasks. The task resolution rate reached 62.1% compared to 57.2% for Claude Opus 4.8, according to the SWE-Atlas QnA benchmark. The performance improvement is attributed to the division of labor and negotiation among agents, highlighting the benefits of multi-agent orchestration over single-agent systems. This advancement could influence the competitive landscape of AI models, with markets estimating an 83.5% probability of Anthropic leading the best AI model rankings by September 2026.
Source: Read the original article

