AI agents capable of using computers have reached a score of 85% on the OSWorld benchmark, compared to just 12% two years ago. In June 2026, Anthropic models dominate the leaderboard: Claude Mythos Preview (85.4%), Fable 5 (85.0%) and Opus 4.8 (83.4%), all surpassing the human baseline of 72%. In December 2025, Simular’s Agent S3 became the first system to cross this human threshold. However, on the harder OSWorld 2.0, which features longer and more complex tasks averaging 1.6 human hours to complete, the best agents complete only 20.6% of assignments. On annotated versions of the tests, they also require between 1.4 and 2.7 times more steps than a human to complete the same tasks.
Source: Read the original article

