The Chinese AI company Z.ai released a new open language model, GLM-5.3-Flash, on August 26, 2026. The weights are available for free under the MIT license on Hugging Face, and the model processes up to one million tokens of context simultaneously. In contrast, the larger sibling model GLM-5.3, which has notably strong cyberattack capabilities, remains locked under security review.
Open weights lower the entry barrier for self-hosters
GLM-5.3-Flash is a mixture-of-experts model with 320 billion parameters, of which only 18 billion are active per request – keeping computing costs and response times low. The model processes text, images, and video natively in a single context window of around one million tokens; this roughly estimates to about 1500 pages of text in a single pass. Those who self-host the weights can download them for free from Hugging Face and run them with common inference tools like vLLM or SGLang.
According to company information, accessing the Z.ai API incurs costs of $0.15 per million input tokens, $0.03 for cached inputs, and $0.50 per million output tokens. Subscribers to the GLM Coding Plan, which starts at $18 per month and is used through the development environment ZCode, receive three times the usage quota for GLM-5.3-Flash compared to the larger GLM-5.3. No restrictions for Germany or the EU are known; access operates as with the predecessor models through the regular Z.ai platform.
Ox Alpha ran undetected on Chinese chips for a week
Starting from August 20, 2026, an unnamed model under the codename “Ox Alpha” appeared on the platforms OpenRouter and OpenCode and quickly became the most used model in the charts within a week. Only with the official announcement did Z.ai confirm that “Ox Alpha” was GLM-5.3-Flash. According to the company, all traffic was fully routed through AI chips manufactured in China, without any Nvidia hardware.
For delivery, Z.ai developed its own serving engine based on SGLang, which, according to the company, triples the response speed compared to standard solutions. The inconspicuous test run under a false name allowed Z.ai to test load and user behavior under real conditions before the brand and price were finalized – a practice that has been rare among Chinese providers so far.
In its own comparative tests, GLM-5.3-Flash performs closely to Anthropic’s top model Claude Opus 4.8: In the Terminal-Bench 2.1, it scores 84.3 out of 100 possible points, while Opus 4.8 scores 85.0. In the internal programming comparison Z.ai Code Bench, both are nearly equal at 29.0 to 29.5 points, independently unverified. Compared to its predecessor GLM-5.2, the score in the coding benchmark DeepSWE v1.1 rises from 46.2 to 63.4 points.
The cyber-capable sibling model remains locked
Unlike GLM-5.3-Flash, the flagship GLM-5.3 introduced in August is still not openly available. The model, which has around 743 billion parameters, achieved 84.5 percent in the security benchmark CyberGym, according to Z.ai, thus reaching a peak ahead of Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol. The ability to independently chain attack steps grew more significantly during training than the company intended.
Therefore, Z.ai announced an additional security review in mid-August and postponed the release of the open weights to the end of August 2026. An analysis by the British AI Security Institute had already shown in July that open models like GLM-5.2 had reduced their lag in attack capabilities compared to top models to four to seven months – a trend that GLM-5.3 claims to continue. For security teams, the staggered rollout means they can already self-host and test the lower-risk GLM-5.3-Flash today, while the potentially more dangerous variant remains accessible only through the controlled API for now.
It will be crucial whether Z.ai actually adheres to the announced release of the GLM-5.3 weights at the end of August or extends the security review again. So far, no other Chinese provider of open models has handled security concerns and open weights as consistently separated as Z.ai with its dual-track GLM-5.3 family – a practice that could set a precedent for future models with similarly striking capabilities.


