Artificial Analysis released a major update to its Coding Agent Index, incorporating reward hacking corrections from Terminal-Bench v2.1. The benchmark now assigns a zero score to any attempt achieving task completion through methods not aligned with intended objectives. Twenty-eight of the 89 tasks in the benchmark were fixed, affecting roughly one-third of the entire suite. GPT-5.6 Sol leads the leaderboard with a score of 89.5%, followed by Claude Opus 5 at 89.1% and Grok 4.6 at 88.4%. The Terminal-Bench v2.1 leaderboard only accepts results run by the benchmark’s own maintainers, disallowing external submissions to ensure a reliable assessment environment.
Source: Read the original article

