Nvidia’s $20B Groq bet goes live this year with new AI racks

1 hour ago 28

Nvidia just turned a $20 billion licensing deal into working hardware. The company confirmed that its Groq 3 LPX low-latency AI inference system has entered full production and will be operational before the end of the year, marking one of the fastest deal-to-deployment timelines in recent AI infrastructure history.

The system, which packs 256 language processing units into a liquid-cooled design, is the first tangible product of Nvidia’s blockbuster licensing agreement with inference chip startup Groq, signed on Christmas Eve 2025. Nebius will be the first company to deploy it through its Token Factory platform.

How a licensing deal became a production rack in eight months

The deal that produced this hardware was unusual by semiconductor industry standards. Rather than acquiring Groq outright, Nvidia paid $20 billion for a non-exclusive license to Groq’s LPU (language processing unit) designs, selectively hired key personnel including Groq founder and CEO Jonathan Ross, and left Groq standing as an independent company. Simon Edwards now leads Groq, which operates 13 data centers globally.

The Groq 3 LPX sits within Nvidia’s Vera Rubin platform, the company’s broader next-generation AI infrastructure stack. The efficiency numbers Nvidia is touting are striking: up to 35 times the throughput per megawatt compared to conventional setups.

Groq’s fundraising momentum and Nvidia’s strategic positioning

Groq raised $650 million in June 2026, followed by an additional $350 million in August 2026, bringing its valuation to $3.5 billion. Nvidia participated in the August round, deepening its financial ties to a company it already pays billions to license from.

Groq operates as an Nvidia Cloud Partner that can integrate Nvidia systems alongside its proprietary offerings.

Nebius and the first real-world deployment

The choice of Nebius as the inaugural deployment partner is notable. The company’s Token Factory platform is designed for large-scale inference workloads, making it a natural testing ground for hardware that claims to dramatically improve inference efficiency.

Nvidia confirmed full production status on August 24, 2026, with deployment planned before year-end.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article