Security

Z.ai launches GLM-5.3-Flash – Cyber model remains locked

3 min read

TL;DR Too Long; Didn’t read

On August 26, 2026, Z.ai released a new open language model GLM-5.3-Flash. The weights are available for free under the MIT license on Hugging Face, and the model processes up to one million tokens of context simultaneously. In contrast, the larger sister model GLM-5.3, which has notably strong cyber attack capabilities, remains locked under security review.

A small padlock with a Z.ai logo sticker stands open, while a larger padlock next to it stays shut with a red warning symbol. Image generated with GPT Image 2

Key takeaways

  • GLM-5.3-Flash uses 320 billion parameters but activates only 18 billion per request.
  • Before the official launch, the model ran unnoticed for about a week as 'Ox Alpha' on OpenRouter and OpenCode.
  • The inference runs entirely on Chinese AI chips, according to Z.ai, with three times faster delivery due to its own SGLang adaptation.
  • API prices are $0.15 per million input tokens and $0.50 per million output tokens.
  • The larger GLM-5.3 with cyber attack capabilities remains under wraps until the end of August, according to the announcement.
  • GLM-Coding-Plan subscribers receive three times the usage quota for GLM-5.3-Flash compared to GLM-5.3.

The Chinese AI company Z.ai released a new open language model, GLM-5.3-Flash, on August 26, 2026. The weights are available for free under the MIT license on Hugging Face, and the model processes up to one million tokens of context simultaneously. In contrast, the larger sibling model GLM-5.3, which has notably strong cyberattack capabilities, remains locked under security review.

Open weights lower the entry barrier for self-hosters

GLM-5.3-Flash is a mixture-of-experts model with 320 billion parameters, of which only 18 billion are active per request – keeping computing costs and response times low. The model processes text, images, and video natively in a single context window of around one million tokens; this roughly estimates to about 1500 pages of text in a single pass. Those who self-host the weights can download them for free from Hugging Face and run them with common inference tools like vLLM or SGLang.

According to company information, accessing the Z.ai API incurs costs of $0.15 per million input tokens, $0.03 for cached inputs, and $0.50 per million output tokens. Subscribers to the GLM Coding Plan, which starts at $18 per month and is used through the development environment ZCode, receive three times the usage quota for GLM-5.3-Flash compared to the larger GLM-5.3. No restrictions for Germany or the EU are known; access operates as with the predecessor models through the regular Z.ai platform.

Ox Alpha ran undetected on Chinese chips for a week

Starting from August 20, 2026, an unnamed model under the codename “Ox Alpha” appeared on the platforms OpenRouter and OpenCode and quickly became the most used model in the charts within a week. Only with the official announcement did Z.ai confirm that “Ox Alpha” was GLM-5.3-Flash. According to the company, all traffic was fully routed through AI chips manufactured in China, without any Nvidia hardware.

For delivery, Z.ai developed its own serving engine based on SGLang, which, according to the company, triples the response speed compared to standard solutions. The inconspicuous test run under a false name allowed Z.ai to test load and user behavior under real conditions before the brand and price were finalized – a practice that has been rare among Chinese providers so far.

In its own comparative tests, GLM-5.3-Flash performs closely to Anthropic’s top model Claude Opus 4.8: In the Terminal-Bench 2.1, it scores 84.3 out of 100 possible points, while Opus 4.8 scores 85.0. In the internal programming comparison Z.ai Code Bench, both are nearly equal at 29.0 to 29.5 points, independently unverified. Compared to its predecessor GLM-5.2, the score in the coding benchmark DeepSWE v1.1 rises from 46.2 to 63.4 points.

The cyber-capable sibling model remains locked

Unlike GLM-5.3-Flash, the flagship GLM-5.3 introduced in August is still not openly available. The model, which has around 743 billion parameters, achieved 84.5 percent in the security benchmark CyberGym, according to Z.ai, thus reaching a peak ahead of Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol. The ability to independently chain attack steps grew more significantly during training than the company intended.

Therefore, Z.ai announced an additional security review in mid-August and postponed the release of the open weights to the end of August 2026. An analysis by the British AI Security Institute had already shown in July that open models like GLM-5.2 had reduced their lag in attack capabilities compared to top models to four to seven months – a trend that GLM-5.3 claims to continue. For security teams, the staggered rollout means they can already self-host and test the lower-risk GLM-5.3-Flash today, while the potentially more dangerous variant remains accessible only through the controlled API for now.

It will be crucial whether Z.ai actually adheres to the announced release of the GLM-5.3 weights at the end of August or extends the security review again. So far, no other Chinese provider of open models has handled security concerns and open weights as consistently separated as Z.ai with its dual-track GLM-5.3 family – a practice that could set a precedent for future models with similarly striking capabilities.

Frequently asked questions

Is GLM-5.3-Flash free to use?

The model weights are available for free under the MIT license on Hugging Face for self-hosting. Those using the API pay per token, with the GLM Coding Plan starting at $18 per month.

What distinguishes GLM-5.3-Flash from the withheld GLM-5.3?

GLM-5.3-Flash is significantly smaller with 320 billion parameters compared to the 743 billion parameter GLM-5.3 and was openly released from the start, without the notably high cyber benchmark score of the larger model.

Is GLM-5.3-Flash usable in Germany or the EU?

No regional lock is known. Access is through the regular Z.ai platform or by self-hosting the open weights, as with the predecessor models.

When are the weights of the large GLM-5.3 expected to be released?

According to an announcement from mid-August 2026, they are expected by the end of August, after completing an additional security review.

How does GLM-5.3-Flash compare to Claude Opus 4.8?

According to Z.ai's own benchmarks, it is close, scoring about 84.3 to 85.0 points in the Terminal-Bench 2.1. The values come from the manufacturer and are independently unverified.

Sources (3)
  1. Z.ai: GLM-5.3-Flash – Frontier Intelligence, Flash Cost
  2. MarkTechPost: Z.ai Releases GLM-5.3-Flash
  3. Z.ai: GLM-5.3 – Frontier Coding with Emergent Cyber Capabilities

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog