Microsoft study reveals long agent runs expose serious reliability issues

Share

A Microsoft Research study using the DELEGATE-52 benchmark shows that frontier AI models suffer significant reliability degradation when facing long, multi-step workflows. Average document fidelity loss reaches approximately 25% after just 20 delegated iterations, climbing to roughly 50% across all models and domains tested. Catastrophic corruption, defined as fidelity dropping to 80% or lower, affects over 80% of model-domain combinations, including Gemini 3.1 Pro, Claude 4.6 Opus and GPT 5.4. Agentic tool usage further worsens the issue, adding an additional 6% degradation. The study reveals that existing safeguards like verification and orchestration fail to address this fundamental reliability challenge in long-horizon delegation.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles