Epoch AI, a nonprofit research institute, launched two puzzle benchmarks designed to test AI reasoning capabilities: Mystery Game Puzzles and Chess Puzzles. Each benchmark contains 100 programmatically generated puzzles to prevent training data contamination. On Mystery Game Puzzles, the top score achieved by any model is 59%, while open-weight models reach a maximum of 38%. On Chess Puzzles, scores progressed from 37% (GPT-5 in December 2025) to 54% currently, showing meaningful improvement over a short period. The gap between frontier closed-source models and open-weight models on reasoning tasks represents a significant competitive disadvantage for decentralized AI projects.
Source: Read the original article

