Nvidia is handing out AI models like free samples at Costco. The logic is the same, too: once you try the product, you’ll come back to buy the hardware that runs it best.
The company’s latest release, Nemotron 3.5 Lightning, is a 30-billion-parameter model that launched on August 11, 2026. It uses a Mixture of Experts (MoE) architecture with roughly 3 billion active parameters, meaning it can run on a single GPU while supporting context windows up to 1 million tokens. That’s a serious amount of capability for a model you can download from Hugging Face without paying a dime.
The razor-and-blades playbook, supercharged
Nvidia now provides free hosted inference for over 100 AI models through its build.nvidia.com APIs. Developers can access these models without swiping a credit card, test them against their own workloads, and build applications on top of them.
Earlier in 2026, Nvidia dropped the Nemotron 3 Ultra, a 550-billion-parameter behemoth, alongside Dynamo 1.0, software that reportedly enhances GPU performance by up to 7x.
Why open models matter for the GPU business
The AI industry has largely split into two camps. On one side, you have companies like OpenAI and Anthropic building closed, proprietary models behind API paywalls. On the other, you have Meta with its Llama series and now Nvidia pushing open-weight alternatives that anyone can download, modify, and deploy.
The models are released under permissive licenses, covering entire families like Nemotron and Cosmos. They’re optimized for Nvidia-specific hardware formats like NVFP4, a quantization format designed for Nvidia chips.
In August 2026, Nvidia also introduced NeMo Switchyard, a tool that intelligently routes tasks across different models to reduce inference costs.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
21








English (US) ·