xAI’s Grok 4.6 just handed the AI benchmark charts a reason to pay attention. Launched in early August 2026, the model scored 1,753 Elo on the GDPval-AA v2 leaderboard, placing third overall and cementing itself as the highest-ranked model from any company not named Anthropic.
That gap from its predecessor is not trivial. Grok 4.5 was sitting in the mid-1500 Elo range before this update, meaning Grok 4.6 added roughly 200 Elo points in a single generational jump.
What GDPval actually measures
OpenAI introduced the GDPval benchmark in September 2025 with a specific goal: stop measuring AI against academic puzzles and start measuring it against the kind of work that actually generates economic value.
The benchmark evaluates models on tasks developed by professionals with an average of 14 years of experience across various industries. The Elo system used here works the same way it does in chess: models compete head-to-head on identical tasks, and ratings shift based on wins and losses.
Artificial Analysis runs the evaluations and maintains the leaderboard, applying the Elo methodology to keep comparisons consistent across model versions and providers. Grok 4.6’s score of 1,753 carries a margin of plus or minus 21 Elo points.
How Grok 4.6 got here
Grok 4.6 is built on the same 1.5 trillion-parameter V9 foundation as Grok 4.5, so the score improvement did not come from simply throwing more compute at the problem. xAI attributes the gains to refined supervised fine-tuning and updated reinforcement learning approaches applied during post-training.
Grok 4.6 sits behind two Claude Opus 5 variants on the leaderboard, which scored 1,849 and 1,817 Elo respectively. The gap between Grok 4.6 and the second Claude Opus 5 variant is about 64 Elo points.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

4 hours ago
13









English (US) ·