Four high-profile artificial intelligence (AI) labs released frontier language models within a three-week stretch in July 2026, and the gap between what these systems finish without human intervention and what came before has narrowed sharply.
Key Takeaways
- xAI shipped Grok 4.5 on July 8 at $2/$6 per million tokens, undercutting Opus 4.8 pricing by over 60%.
- Anthropic released Claude Opus 5 on July 24 as the new default model on Claude Max subscriptions.
- Moonshot AI published Kimi K3, its largest model at 2.8 trillion parameters.
Grok 4.5 from xAI, Claude Opus 5 from Anthropic, the GPT-5.6 family from OpenAI, and Kimi K3 from Moonshot AI each target the same problem: getting an AI system to carry a multi-hour task, from research to coding to structured reporting, without losing track of the plan.
Grok 4.5 Trains on Real Developer Sessions
xAI released Grok 4.5 on July 8. The model runs on a 1.5 trillion parameter base and was trained in part on real usage data from Cursor, the coding platform SpaceXAI acquired earlier this year. Pricing sits at $2 per million input tokens and $6 per million output tokens, with a 500,000 token context window.
Elon discussing the release of Grok 4.5 on July 8, 2026.Elon Musk described Grok 4.5 as an Opus-class model that runs faster and at lower cost. Independent trackers place it fourth on the Artificial Analysis Intelligence Index, ahead of every open weight model, at pricing more than 60% below Claude Opus 4.8 and GPT-5.5. On Terminal-Bench 2.1, a test that scores how well a model completes command-line engineering tasks, Grok 4.5 scored 83.3%.
The bigger shift is token efficiency. xAI says the model needs roughly a fifth of the output tokens Opus 4.8 required for comparable tasks, which lowers the cost of running long agent sessions.
Claude Opus 5 Holds the Line on Price
Anthropic released Claude Opus 5 on July 24, positioning it as a model that reaches close to the performance of Claude Fable 5, the company’s most capable release, at half of Fable’s $10 input and $50 output pricing. Opus 5 itself carries the same $5 input and $25 output pricing as its predecessor, Opus 4.8.
Claude Opus 5 announcement via X.The model ships with a 1 million token context window, 128,000 max output tokens, and an adjustable reasoning effort setting that ranges from low to a new xhigh mode. Anthropic says Opus 5 sets new marks on Frontier-Bench and GDPval-AA, two coding and knowledge work evaluations, though it trails the restricted Claude Mythos 5 model on cybersecurity tasks. Opus 5 is now the default model on Claude Max and the strongest option on Claude Pro.
OpenAI Splits GPT-5.6 Into Three Tiers
OpenAI moved GPT-5.6 to general availability on July 9 after a two-week preview limited to roughly 20 organizations vetted by the U.S. government, following an executive order tied to frontier model safety review. The family ships as three tiers: Sol, the flagship, priced at $5 input and $30 output per million tokens; Terra, a mid-tier model at $2.50 and $15 that OpenAI says matches GPT-5.5 at half the cost; and Luna, a fast, low-cost tier at $1 and $6.
ChatGPT 5.6 Sol announcement via X.OpenAI reports Sol leads the Artificial Analysis Coding Agent Index and hits 88.8% on Terminal-Bench 2.1, rising to 91.9% when the model runs four sub-agents in parallel under its new ultra mode. All three tiers carry OpenAI’s highest internal risk rating for cyber and biological misuse potential, which triggered added review during the government-gated preview.
Kimi K3 Pushes Open Weight Models Toward the Frontier
Moonshot AI released Kimi K3 on July 16, a 2.8 trillion parameter mixture of experts model with 896 total experts and 16 active per task. The model carries a 1 million token context window and native multimodal input, with API pricing at $3 input and $15 output per million tokens. Moonshot has committed to publishing full open weights by July 27.
Kimi K3 announcement via X.K3 is the largest open-weight model released to date, roughly 75% bigger than the previous largest widely used open model. Independent trackers place K3 fourth among current frontier systems, behind Claude Fable 5 and GPT-5.6 Sol but ahead of Claude Opus 4.8.
Why the Gap With Elder Models Matters
Context windows below 200,000 tokens once forced developers to break large codebases or research packets into fragments. Every model in this group now runs at 500,000 tokens or beyond, with three of the four at 1 million, letting a single session hold a full repository or a stack of primary source documents.
Agent reliability has moved as well. Systems from the 2023 and 2024 period often lost their plan after a handful of tool calls. The models released this month are built to sustain dozens of coordinated steps and recover when a tool call returns bad data instead of stalling out.
For developers, the practical effect is a lower cost per finished task rather than a higher ceiling on any single benchmark. Effort controls in Opus 5, tiered pricing in GPT-5.6, and the efficiency claims behind Grok 4.5 point toward the same goal: letting teams choose how much compute a task deserves instead of paying flagship prices for every request.
For enterprises and policymakers, the government-gated rollout of GPT-5.6 signals that oversight is now built into release schedules for the largest models, not added after the fact. Anthropic’s brief, government-directed suspension of Fable and Mythos access in June was an earlier version of the same pattern.
XRP has spent much of 2026 at the center of speculation, fueled by institutional adoption, regulatory developments, and renewed interest…
$1.32 to $3.75: Nine Top AI Models Predict XRP’s 2026 Finish, One Breaks Away
XRP has spent much of 2026 at the center of speculation, fueled by institutional adoption, regulatory developments, and renewed interest…
$1.32 to $3.75: Nine Top AI Models Predict XRP’s 2026 Finish, One Breaks Away
XRP has spent much of 2026 at the center of speculation, fueled by institutional adoption, regulatory developments, and renewed interest…

1 hour ago
24









English (US) ·