If you’ve ever watched someone try to figure out the rules of a board game mid-play, you have a rough idea of what Epoch AI is now asking artificial intelligence to do. The nonprofit research institute has rolled out a pair of game-based benchmarks, Mystery Game Puzzles and Chess Puzzles, designed to stress-test the reasoning abilities that AI companies love to brag about.
The early results are humbling. The top score on the Mystery Game Puzzles sits at 59%, and open-weight models max out at just 38%.
What the benchmarks actually measure
Each benchmark consists of 100 programmatically generated puzzles. The Chess Puzzles are relatively straightforward in concept: given a board position, find the best move. The Mystery Game Puzzles are something different entirely.
In the mystery variant, the AI doesn’t even know which game it’s playing. The game’s identity is deliberately obscured, which means the model can’t fall back on memorized patterns or training data shortcuts.
This design choice is intentional. Traditional AI benchmarks have a contamination problem. Models train on massive internet datasets that often include the very tests they’re evaluated on. By generating puzzles programmatically and hiding the game format, Epoch AI strips away the safety net of pattern recognition and forces models to demonstrate genuine spatial reasoning and planning.
Responses are scored through normalized move notation against an established answer key. No agent scaffolding, no chain-of-thought prompting tricks, no tool use.
The scorecard so far
On Mystery Game Puzzles, the ceiling is 59%. That’s the best any model has managed. Open-weight models top out at 38%. The gap between frontier closed-source models and their open-weight counterparts is stark.
Chess Puzzles tell a slightly more optimistic story. GPT-5 scored 37% when the benchmark launched in December 2025. Newer models have since pushed that figure to 54%, showing meaningful improvement over a relatively short window. But even 54% on chess puzzles highlights an important distinction: being good at playing chess with a search engine is very different from reasoning through novel positions without one.
Neither benchmark appears close to saturation.
Why this matters beyond AI research
Epoch AI maintains a broader benchmarking hub and the Epoch Capabilities Index, which aggregates evaluations across mathematics, coding, and gameplay. The ECI is designed to provide quick estimates of model capability without relying on any single test. A related benchmark called EBR-bench, based on the game Earthborne Rangers, showed limited AI improvement in learning from repeated engagements as of July 2026, reinforcing the pattern that interactive reasoning remains a stubborn frontier.
The improvement from 37% to 54% on chess puzzles over a period of months shows that progress hasn’t stopped. But the mystery game results suggest that generalized reasoning is advancing more slowly than the marketing materials imply.
At 38% versus 59%, open-weight models are significantly behind closed-source models on reasoning tasks. For decentralized AI projects that rely on open models, this performance delta represents a real competitive disadvantage that won’t be solved by simply adding more GPUs to the network.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
20








English (US) ·