Z.ai's GLM-5.3 Edges Past Mythos and GPT-5.6 on a Cybersecurity Benchmark
China's Zhipu shipped a new flagship without touching the base model. Every gain came from post-training. Its 84.5 on CyberGym lands above Anthropic's Mythos and OpenAI's GPT-5.6 Sol, but the weights are held back two weeks.

Zhipu, the Chinese lab that ships under the Z.ai brand, released GLM-5.3 on August 14.
The odd part is that the base model did not change at all. It is the same 743-billion-parameter MoE architecture as GLM-5.2.
So what did change?
Post-training is the stage after the main training run that tunes a model for specific work.
Z.ai claims coding capability improved by roughly 50%, and that the model now tops its in-house Code Bench among open-weights models.
That is a meaningful technical claim in itself: in 2026 you do not necessarily need a new base model to make a step change. You can spend the budget on the stage after it.
The security result is the real story
The number drawing attention is on CyberGym, a benchmark that measures whether a model can identify and validate security flaws from source code.
GLM-5.3 scored 84.5%. For comparison, Anthropic's Mythos scored 83.8% and OpenAI's GPT-5.6 Sol 83.6%.
These figures come from Zhipu and have not been independently verified, but the margins are narrow enough to carry their own message.
Another number frames it better, and it reads differently next to Anthropic's 45-agent swarm experiment: since GLM-5.2, Z.ai's models have collectively identified 2,436 vulnerabilities across 269 open-source projects, 1,097 of them rated critical or high severity.
The weights are held back
Unlike previous releases in this family, the weights did not ship alongside the model. Z.ai says they arrive in two weeks.
It is the first such delay in the GLM-5 line, and given that vulnerability discovery is the model's headline strength, the reasoning is not especially mysterious.
A model that can find a security bug is the same model that can exploit one. 🔍




