A study presented by Microsoft highlights significant limitations of artificial intelligence systems in long-term decision-making tasks. AI models achieved only 27% of human performance when tasked with interconnected decisions over a simulated year. The best configuration tested, Qwen3.7-Max with Hermes, remains significantly below human benchmarks, underscoring the gap in long-horizon reliability. These findings could impact Anthropic’s position in the AI model race.
Source: Read the original article
Disclaimer: this content is for information purposes only and is not financial advice. Cryptocurrencies are highly volatile: you may lose all of your capital. Always do your own research. Legal notice

