Microsoft is reportedly evaluating Moonshot AI’s Kimi K3 model for integration into its Copilot service on Azure, a move that would push two of the biggest names in AI, OpenAI and Anthropic, at least partially to the sidelines on one of the world’s largest cloud platforms.
The timing is notable. Kimi K3 was only released on July 16, 2026, and it already has people paying attention for the right reasons.
What Kimi K3 actually is
Kimi K3 is a multimodal model built by Beijing-based Moonshot AI, packing 2.8 trillion parameters and a context window of one million tokens.
The model scored 1,679 points on LMArena’s Frontend Code Arena, landing in first place and beating out the current field of competitors.
Kimi K3 is also priced aggressively. Input tokens cost $3 per million, while output tokens run $15 per million.
Demand hit so hard after launch that Moonshot AI paused new K3 subscriptions due to compute strain. The company plans to release full open weights by July 27, 2026, which would allow enterprises and developers to run the model on their own infrastructure.
Microsoft and Moonshot have history
This would not be Moonshot AI’s first appearance on Azure. Kimi K2 Thinking was integrated into Microsoft Foundry in December 2025, followed by Kimi K2.5 in February 2026.
Microsoft has not made a public statement confirming any plan to replace existing OpenAI or Anthropic model integrations. The current situation is best described as an evaluation phase, which is standard practice before any large-scale model deployment on a platform like Azure.
Why this matters for the AI market
Microsoft has a multi-billion-dollar relationship with OpenAI, so the fact that it is even exploring alternatives tells you something about how the competitive dynamics in foundation model markets are shifting.
The semiconductor market has already felt some of this anxiety. Kimi K3’s release triggered declines in chip stocks, echoing the pattern seen earlier in 2025 when DeepSeek’s efficient model architecture rattled investor confidence in the assumption that AI scaling always requires proportionally more hardware.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

14 hours ago
23









English (US) ·