Epoch AI launches game puzzles benchmark that has AI models stuck at 59%

Share

Epoch AI, a nonprofit research institute, launched two puzzle benchmarks designed to test AI reasoning capabilities: Mystery Game Puzzles and Chess Puzzles. Each benchmark contains 100 programmatically generated puzzles to prevent training data contamination. On Mystery Game Puzzles, the top score achieved by any model is 59%, while open-weight models reach a maximum of 38%. On Chess Puzzles, scores progressed from 37% (GPT-5 in December 2025) to 54% currently, showing meaningful improvement over a short period. The gap between frontier closed-source models and open-weight models on reasoning tasks represents a significant competitive disadvantage for decentralized AI projects.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles