Alibaba’s Qwen 3.8-Max enters enterprise market with 2.4 trillion parameters and aggressive pricing

1 hour ago 23

Alibaba has rolled out Qwen 3.8-Max, a 2.4 trillion parameter AI model previewed on July 19-20, packing roughly 95 billion active parameters and capable of processing context windows of approximately 1 million tokens.

Standard API pricing lands at about $2 per million input tokens and $6 per million output tokens. Preview users get a 10% discount off standard rates during the promotional period.

What Qwen 3.8-Max actually does

The model is multimodal, meaning it processes text, images, and video rather than just written language.

A context window of roughly 1 million tokens means the model can ingest entire codebases, lengthy legal documents, or multi-chapter reports in a single pass.

Alibaba has highlighted two areas where Qwen 3.8-Max is particularly strong: coding tasks and agentic workflows. The latter refers to AI systems that can chain together multiple steps to complete complex tasks autonomously.

The company also plans to release open weights for the Max-class model, continuing a pattern established with earlier versions of the Qwen series. Open weights allow developers and researchers to fine-tune the model for specialized use cases.

Pricing strategy signals a land grab

At $2 per million input tokens, the company is positioning Qwen 3.8-Max as a cost-effective alternative to comparable Western models. There are no confirmed reports of a revenue-sharing model for enterprise users; pricing is based on token usage and subscriptions.

Enterprise integration tools like DingTalk, Alibaba’s workplace collaboration platform, add another layer to the strategy by bundling AI capabilities with existing enterprise software.

Where this fits in the global AI landscape

The 2.4 trillion total parameter count puts Qwen 3.8-Max among the largest publicly acknowledged AI models in the world. The 95 billion active parameter figure indicates Alibaba is using a mixture-of-experts architecture where only a fraction of the model’s total capacity is engaged for any given task.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article