Alibaba has released its AI flagship Qwen3.8-Max for paying customers on August 3, 2026, thus ending the preview phase that has been running since July. According to the company, the language model activates only 95 billion of the total 2.4 trillion parameters per request. Open weights for standalone operation are expected to follow as early as next week.
Alibaba answers the open question about the activation rate
The transition from preview to full version clarifies a central question that beckmann.ai left open in July. How many of the 2.4 trillion parameters Qwen3.8-Max actually activates per request has remained unanswered until now. According to the official blog post from Alibaba Cloud, a Sparse-Mixture-of-Experts architecture is used, where only 95 billion parameters are computed – just under four percent of the total size. This figure determines the actual computing costs and was not known for its predecessor. The model also processes up to one million tokens of context and, in addition to text, also images and videos. This allows Qwen3.8-Max to achieve similar key figures as Kimi K3 from Moonshot AI, which also boasts one million tokens of context and multimodal capabilities.
An additional, somewhat independent indicator is provided by the chatbot ranking Arena.ai: There, Qwen3.8-Max ranks fifth in the text category and second in visual tasks. Unlike Alibaba’s own benchmark tables, the Arena ranking is based on anonymous user ratings and cannot be solely controlled by the manufacturer. However, how robust this interim result is depends on how many ratings are available so far – neither Alibaba nor the Arena operator provided specific information on this.
Benchmarks show lead over Fable 5 and GPT-5.6 Sol
According to Alibaba, Qwen3.8-Max performs better in seven programming and general tests as well as in 36 multimodal benchmarks than Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol. The company positions its model for the first time not just as number two behind the competition, but as equal or superior. Fable 5 is considered particularly powerful: the USA temporarily imposed export restrictions on the model according to a Bloomberg report because it was classified as strategically significant. A comparison with such a classified system further enhances Alibaba’s results – at least rhetorically. However, Alibaba’s own benchmark table is not independently verified; third parties have not yet corroborated the figures.
Thus, Qwen3.8-Max joins a series of Chinese models that have made similar claims in recent weeks – including DeepSeek V4 Flash and Tencent’s Hy3. Each of these models advertises high rankings in specially selected tests, but independent comparative measurements remain the exception. As with the predecessor model Qwen3.7-Max, it has been shown in the past that even published best values turned out to be more moderate after external testing.
Access currently only through the paid API
Developers can initially access Qwen3.8-Max exclusively through the paid API of Alibaba Cloud Model Studio, which is also accessible from Germany and the EU via the international endpoint in Singapore. According to a pricing overview from the intermediary OpenRouter, Alibaba charges two US dollars per million input tokens and six US dollars per million output tokens – more than the predecessor Qwen3.7-Max, which was priced at 1.25 and 3.75 US dollars respectively. A free quota for new accounts exists, but it only covers a limited testing period.
The open model weights for local operation without cloud connection are expected to appear on Hugging Face and ModelScope next week, although the company did not specify an exact date. Only with the open weights can the model be operated independently of Alibaba’s own infrastructure and tested outside of the API billing. For companies in Germany, this means for now: those who want to use Qwen3.8-Max today will pay based on usage through the cloud API; on-premise use will only be possible after the weight release.
It will be crucial whether external laboratories confirm the best values claimed by Alibaba against Fable 5 and GPT-5.6 Sol once the open weights are available. So far, every ranking between the Chinese top models and their US competitors has come solely from the manufacturers themselves. The next test will therefore fall on the coming week when Qwen3.8-Max can be tested for the first time outside of Alibaba’s own cloud environment.


