Epoch classifies first AI solution as major advance in FrontierMath benchmark

1 hour ago 25

When Epoch AI launched Tier 4 of its FrontierMath benchmark in July 2025, the best AI models on the planet could solve roughly 5% of its problems. Fourteen months later, OpenAI’s GPT-6 Astra has cracked nearly all of them, earning the first-ever “Major Advance” classification from Epoch for an AI-generated mathematical solution.

The model achieved between 97.6% and 98% accuracy across all 43 problems in Tier 4, the benchmark’s most difficult category.

The problem that was supposed to be unsolvable

The last Tier 4 problem to fall was authored by combinatorialist Jay Pantone. It was specifically designed to resist the kinds of shortcuts that earlier AI models had exploited to inflate their scores on other FrontierMath problems.

GPT-6 Astra solved it anyway, which is precisely why Epoch elevated the result to “Major Advance” status. The classification isn’t just about getting the right answer. It reflects Epoch’s assessment that the solution demonstrates genuinely novel mathematical reasoning capability rather than pattern-matching tricks.

This is the first time any AI system has received this designation from Epoch’s FrontierMath program.

What FrontierMath actually measures

FrontierMath isn’t your standard math test. Epoch AI designed it as a collection of research-grade mathematical problems spanning multiple tiers of difficulty, with Tier 4 sitting at the top as problems approaching the frontier of human mathematical research.

After corrections made in June 2026, the benchmark now comprises around 295 problems across tiers 1 through 3, plus the 43 problems in Tier 4.

The rapid saturation of Tier 4 has already prompted Epoch to shift its focus. The organization is now building out more challenging collections, including what it calls the FrontierMath Erdős set, named after the legendarily prolific mathematician Paul Erdős.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article