AI-Models

Alibaba releases flagship model Qwen3.8-Max

4 min read

TL;DR Too Long; Didn’t read

Alibaba unlocked Qwen3.8-Max on August 3, 2026, via the cloud platform Model Studio. The system activates only 95 out of 2.4 trillion parameters per request and processes up to one million tokens of context, according to the manufacturer. According to Alibaba, it outperforms Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol in several tests. Open weights for local use are expected to follow next week.

A server tower with an attached Alibaba Cloud logo, from which only a handful of bright points shine out of its huge grid of thousands of small light points, next to an upward-pointing bar chart Image generated with GPT Image 2
Part of the dossier: Qwen3.8-Max: From preview to full release →

Key takeaways

  • Only 95 of the 2.4 trillion parameters are active per request – the previously unanswered metric from the preview.
  • The context window processes up to one million tokens, similar to competitor Kimi K3.
  • Alibaba Cloud charges 2 US dollars per million input tokens and 6 US dollars per million output tokens.
  • In the Arena.ai ranking, the model ranks fifth in text and second in the vision area.
  • Alibaba announces open model weights for Hugging Face and ModelScope for the coming week.
  • Fable 5, the comparison model from Anthropic, was temporarily subject to export restrictions in the USA.

Alibaba has released its AI flagship Qwen3.8-Max for paying customers on August 3, 2026, thus ending the preview phase that has been running since July. According to the company, the language model activates only 95 billion of the total 2.4 trillion parameters per request. Open weights for standalone operation are expected to follow as early as next week.

Alibaba answers the open question about the activation rate

The transition from preview to full version clarifies a central question that beckmann.ai left open in July. How many of the 2.4 trillion parameters Qwen3.8-Max actually activates per request has remained unanswered until now. According to the official blog post from Alibaba Cloud, a Sparse-Mixture-of-Experts architecture is used, where only 95 billion parameters are computed – just under four percent of the total size. This figure determines the actual computing costs and was not known for its predecessor. The model also processes up to one million tokens of context and, in addition to text, also images and videos. This allows Qwen3.8-Max to achieve similar key figures as Kimi K3 from Moonshot AI, which also boasts one million tokens of context and multimodal capabilities.

An additional, somewhat independent indicator is provided by the chatbot ranking Arena.ai: There, Qwen3.8-Max ranks fifth in the text category and second in visual tasks. Unlike Alibaba’s own benchmark tables, the Arena ranking is based on anonymous user ratings and cannot be solely controlled by the manufacturer. However, how robust this interim result is depends on how many ratings are available so far – neither Alibaba nor the Arena operator provided specific information on this.

Benchmarks show lead over Fable 5 and GPT-5.6 Sol

According to Alibaba, Qwen3.8-Max performs better in seven programming and general tests as well as in 36 multimodal benchmarks than Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol. The company positions its model for the first time not just as number two behind the competition, but as equal or superior. Fable 5 is considered particularly powerful: the USA temporarily imposed export restrictions on the model according to a Bloomberg report because it was classified as strategically significant. A comparison with such a classified system further enhances Alibaba’s results – at least rhetorically. However, Alibaba’s own benchmark table is not independently verified; third parties have not yet corroborated the figures.

Thus, Qwen3.8-Max joins a series of Chinese models that have made similar claims in recent weeks – including DeepSeek V4 Flash and Tencent’s Hy3. Each of these models advertises high rankings in specially selected tests, but independent comparative measurements remain the exception. As with the predecessor model Qwen3.7-Max, it has been shown in the past that even published best values turned out to be more moderate after external testing.

Access currently only through the paid API

Developers can initially access Qwen3.8-Max exclusively through the paid API of Alibaba Cloud Model Studio, which is also accessible from Germany and the EU via the international endpoint in Singapore. According to a pricing overview from the intermediary OpenRouter, Alibaba charges two US dollars per million input tokens and six US dollars per million output tokens – more than the predecessor Qwen3.7-Max, which was priced at 1.25 and 3.75 US dollars respectively. A free quota for new accounts exists, but it only covers a limited testing period.

The open model weights for local operation without cloud connection are expected to appear on Hugging Face and ModelScope next week, although the company did not specify an exact date. Only with the open weights can the model be operated independently of Alibaba’s own infrastructure and tested outside of the API billing. For companies in Germany, this means for now: those who want to use Qwen3.8-Max today will pay based on usage through the cloud API; on-premise use will only be possible after the weight release.

It will be crucial whether external laboratories confirm the best values claimed by Alibaba against Fable 5 and GPT-5.6 Sol once the open weights are available. So far, every ranking between the Chinese top models and their US competitors has come solely from the manufacturers themselves. The next test will therefore fall on the coming week when Qwen3.8-Max can be tested for the first time outside of Alibaba’s own cloud environment.

Frequently asked questions

What is the cost of access to Qwen3.8-Max?

According to third-party information, Alibaba charges 2 US dollars per million input tokens and 6 US dollars per million output tokens via the API of Alibaba Cloud Model Studio. New accounts receive a limited free trial quota.

When will the open weights of Qwen3.8-Max be released?

Alibaba announces the release on Hugging Face and ModelScope for the coming week, but the company has not yet provided a specific date.

What has changed since the preview in July?

The unpaid preview has become a paid full version. The activation rate of 95 billion parameters and benchmark comparisons to Fable 5 and GPT-5.6 Sol, which were missing in July, are now known.

Is Qwen3.8-Max better than Kimi K3 from Moonshot AI?

An independent direct comparison is not yet available. Both models process up to one million tokens of context, Kimi K3 has more total parameters with 2.8 trillion and is already openly available.

Can companies in Germany already use Qwen3.8-Max?

Yes, via the international API endpoint in Singapore, billed by usage. Operation on their own infrastructure is only possible once the open weights are released.

Sources (4)
  1. Alibaba Cloud Community: Alibaba Unveils Qwen3.8-Max
  2. Bloomberg: Alibaba Adds to China AI Breakthroughs With New Qwen Model
  3. InfoWorld: Alibaba takes aim at OpenAI and Anthropic with Qwen3.8-Max launch
  4. OpenRouter: Qwen3.8-Max pricing overview

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog