Artificial Analysis updates Coding Agent Index with reward hacking corrections

Share

Artificial Analysis released a major update to its Coding Agent Index, incorporating reward hacking corrections from Terminal-Bench v2.1. The benchmark now assigns a zero score to any attempt achieving task completion through methods not aligned with intended objectives. Twenty-eight of the 89 tasks in the benchmark were fixed, affecting roughly one-third of the entire suite. GPT-5.6 Sol leads the leaderboard with a score of 89.5%, followed by Claude Opus 5 at 89.1% and Grok 4.6 at 88.4%. The Terminal-Bench v2.1 leaderboard only accepts results run by the benchmark’s own maintainers, disallowing external submissions to ensure a reliable assessment environment.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles