Stanford paper reveals AI teams outperform debate-and-vote models

Share

A Stanford University study introduces the Self-Organizing Agent Teams (SAT) framework, in which groups of AI agents learn to collaborate from a small set of past experiences. This approach achieved an average accuracy of 66.7% across five math and physics benchmarks, significantly outperforming the best individual agent at 48.8%, a compute-matched single agent at 58.7%, and a routing oracle at 59.0%. SAT teams derived their collaborative strategies from just 15 AIME 2024 problems or 25 GPQA Diamond problems, then applied them to entirely new benchmarks. On AIME 2026 specifically, SAT reached 71.2% accuracy, exceeding the routing oracle by 13.4 percentage points. The researchers note that the 0.90 correlation with « demonstrability » indicates the framework works best on problems where correct reasoning can be easily identified in conversation, such as in math and physics.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles