Here’s a fun paradox for the AI industry: the better your models get, the less your scorecard matters. That’s essentially what Cognition CEO Scott Wu is arguing as the company’s flagship AI coding agent, Devin, approaches near-perfect scores on the very benchmarks that once defined the competitive landscape.
Wu’s position is that traditional AI benchmarks are becoming less meaningful because frontier models can now solve essentially any well-defined task thrown at them. When everyone’s acing the test, the test stops telling you anything useful.
From 13% to 90%, and now what
When Devin launched in March 2024, it scored just 13% on SWE-Bench, a widely used benchmark for evaluating AI software engineering capabilities. By mid-2026, that number had climbed to roughly 90% on the original SWE-Bench and approximately 80% on SWE-Bench Pro.
Instead of chasing public benchmark scores, Cognition has developed its own proprietary evaluation called FrontierCode 1.1. The company also uses internal “junior dev” benchmarks designed to simulate real-world coding tasks rather than the kind of cleanly defined problems that traditional benchmarks tend to favor.
Wu has been particularly pointed about metrics like token usage, arguing they can actually mislead evaluations of what an AI system is genuinely producing.
A $26 billion bet on outcomes over scores
Cognition raised $1 billion in May 2026 at a $26 billion valuation. The company was founded in November 2023 by Wu and his co-founders.
In certain contexts, Devin reportedly contributes about 89% of the committed code produced by Cognition’s own engineering team.
Cognition also recently acquired Windsurf, a rival AI coding entity, in a move that bolsters its competitive position as the market for autonomous coding tools heats up.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
26









English (US) ·