Security

GLM-5.3: Z.ai holds back model weights over cyber risk

3 min read

TL;DR Too Long; Didn’t read

Z.ai released the open language model GLM-5.3 on August 14, 2026, which scored 84.5 percent on the security benchmark CyberGym, narrowly ahead of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. According to the company, its ability to plan attacks independently grew far more than intended, so the open model weights won't follow until late August — until then, access is limited to the API and subscription.

A hand holds back a keychain with a Z.ai logo sticker, while one key turns into a digital lockpick breaking open a padlock. Image generated with GPT Image 2

Key takeaways

  • GLM-5.3 outperforms Mythos 5 and GPT-5.6 Sol on the CyberGym security benchmark with 84.5 percent.
  • According to Z.ai, the model's attack-planning ability grew far more than intended during training.
  • The model found 2,436 vulnerabilities in 269 open-source projects, 1,097 of them rated critical or high.
  • Z.ai is delaying the public weight release for additional security review until the end of August 2026.
  • Access starts at $18 a month via the GLM Coding Plan; a per-token API price is still pending.
  • GLM-5.2 and GLM-5.3 share the same roughly 743-billion-parameter base; all gains come from extra training.

The Chinese AI company Z.ai has introduced an open language model, GLM-5.3, which achieved 84.5 percent in the security benchmark CyberGym, narrowly ahead of Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol. According to the company, the ability to independently plan cyberattacks grew significantly more than intended. As a result, Z.ai is delaying the public release of the model weights by several weeks.

GLM-5.3 surpasses Anthropic and OpenAI in attack benchmark

Z.ai builds GLM-5.3 on the same base of roughly 743 billion parameters as its predecessor GLM-5.2 – all improvements come solely from additional training after the actual pre-training, not from a new base model. In CyberGym, a benchmark that measures the independent exploitation of known software vulnerabilities in isolated test environments, the score rises from 77.2 to 84.5 percent. In the tougher ExploitBench, the success rate more than doubles, from 24.4 to 54.4 percent. In the timed ExploitGym benchmark, the model solves 105 tasks within two hours instead of the previous 29, and 130 tasks within six hours instead of 39. GLM-5.3 also improves significantly at programming: on Terminal-Bench 3.0, the score climbs from 4.6 to 28.3 points, and on the coding benchmark DeepSWE v1.1, from 46.2 to 66.9. These numbers come from Z.ai’s own publications and are independently unverified. For users, this mainly means that an open model from China can now independently handle more complex, multi-step programming and security tasks that previously required mostly closed frontier models.

Model develops attack plans without targeted training for it

According to its own statements, Z.ai had specifically added vulnerability-discovery data during the training phase to teach the model to find individual software flaws. But the capability grew beyond what was expected: GLM-5.3 independently chains multiple attack steps into coherent exploitation plans – an ability the company says was not planned and only became apparent during training. In a practical test, the model found 2,436 vulnerabilities across 269 open-source software projects, of which 1,097 were rated high or critical severity – roughly nine vulnerabilities per project examined on average, about four of them severe. Already in July, an analysis by the UK’s AI Security Institute had shown that open models like GLM-5.2 had narrowed their lag in attack capabilities behind top models to four to seven months. GLM-5.3 continues that trend and, based on the available figures, narrows the gap further. For security teams, this means growing time pressure: tools that independently find and chain vulnerabilities shrink the window between the discovery of a flaw and its exploitation even further.

Open weights follow only after additional security review

Z.ai is initially making GLM-5.3 available only in a limited way: the model can be accessed through the GLM Coding Plan, a subscription starting at $18 a month for the entry-level Lite tier, $80 for Pro, and $168 for Max. A per-token price for direct API access had not been announced at the time of publication; for the predecessor GLM-5.2, it was $1.40 per million input tokens and $4.40 per million output tokens. According to the company, both broader API access and the public model weights are expected to follow only after security review and hardening are complete, likely by the end of August 2026 – roughly two weeks after the current launch. No specific restriction for Germany or the EU is known; no official statement on this exists, and access runs the same way as with GLM-5.2, through the regular Z.ai platform. Anyone who wants to use the model right now therefore needs a subscription and cannot fall back on a free, self-hosted version.

What remains open is whether Z.ai will use the extra time before the weight release for an external review by independent security researchers, or whether the delay will run purely internally. A staggered, security-reviewed rollout would be a first among Chinese open-weight providers, which have so far mostly released their models without a separate approval phase.

Frequently asked questions

When will GLM-5.3's open weights be released?

According to the company, likely by the end of August 2026, about two weeks after launch, once additional security reviews are complete.

What does access to GLM-5.3 cost?

The GLM Coding Plan starts at $18 a month for the Lite tier, $80 for Pro, and $168 for Max. A separate per-token API price had not been set as of release.

How does GLM-5.3 compare to Mythos 5 and GPT-5.6 Sol?

On the CyberGym security benchmark, GLM-5.3 scored 84.5 percent, according to Z.ai — narrowly ahead of Mythos 5 at 83.8 percent and GPT-5.6 Sol at 83.6 percent.

Is GLM-5.3 available in Germany or the EU?

No regional restriction is known. Access works the same way as with the predecessor model, through the regular Z.ai platform, and the company has not mentioned any EU-specific limitation.

What technically sets GLM-5.3 apart from GLM-5.2?

Both models use the same base architecture with roughly 743 billion parameters. According to Z.ai, all improvements come solely from additional training after pre-training, not from a new base model.

Sources (3)
  1. Z.ai: GLM-5.3 – Frontier Coding with Emergent Cyber Capabilities
  2. Unite.AI: Z.ai Launches GLM-5.3 With Frontier Coding and a Cyber Capability That Outgrew Its Training
  3. MarkTechPost: Z.ai Ships GLM-5.3 Without Retraining the Base Model

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog