AI ONLINE6 September 2026
The AI News Desk

RelayON THE WIRE

The whole field of AI — read, checked, and explained.
Models & Releases

GLM-5.3 Tops the Open Coding Field on Post-Training Alone — and the Cyber Jump Is Why the Weights Are Waiting

Z.ai's new model reuses GLM-5.2's exact 743B base and improves through post-training alone — the strongest open-weights coder yet. And in a departure for Z.ai, which normally ships its weights within days, it's holding them back for a safety review, because the capability that jumped most was cyber.

Des OkoroBy Des OkoroResearch Correspondent
19 August 2026
Listen to this postread by Relay

When Z.ai shipped GLM-5.2 earlier this summer, it did the thing open-weights labs are supposed to do: it put the weights on Hugging Face, under an MIT licence, more or less the same week. GLM-5.3, which landed on 14 August, tops its predecessor on almost every coding benchmark that matters — and this time the weights are not coming with it. They ship "in stages, following rigorous safety evaluations," roughly two weeks behind the model itself.

The gap between those two decisions is the actual story.

Same base, all post-training

The first surprising thing about GLM-5.3 is what did not change. It reuses the exact 743-billion-parameter Mixture-of-Experts base from GLM-5.2 — no new pre-training, no larger model. Z.ai is blunt about it: "Scaling post-training is all we did for GLM-5.3." The gains came from the recipe — longer task environments, synthesized end-to-end training data, and verifier agents that confirm a task is actually solvable before it is used for reinforcement learning.

The results are not marginal. On DeepSWE, GLM-5.3 goes from GLM-5.2's 46.2 to 66.9. On Terminal-Bench 3.0 it jumps from 4.6 to 28.3 — and it posts the highest GDPval-AA v2 score of anything Z.ai tested. It is, on the coding benchmarks, the strongest open-weights model available.

Two caveats keep that claim honest. "Strongest open model" is not "strongest model" — on Terminal-Bench 3.0 its 28.3 still trails the closed frontier, where Claude Fable 5 (33.7) and GPT-5.6 Sol (34.6) sit ahead of it. And a lab grading its own homework is a lab grading its own homework: these are Z.ai's numbers on Z.ai's chart, and the independent re-runs will matter.

But the headline holds up: you can take a fixed base and, through post-training alone, close a large part of the distance to the frontier. That is a useful, concrete data point about where the cheap gains still are.

The cyber jump is the part to watch

The capability that moved most is the one nobody advertises on a launch chart. On CyberGym, GLM-5.3 scores 84.5% — first place, narrowly ahead of the frontier models behind it at 83.8% and 83.6%. On ExploitBench it more than doubled its predecessor, from 24.4% to 54.4%. In real-world testing Z.ai reports the model surfaced 2,436 vulnerabilities across 269 projects, over a thousand of them medium-to-high severity.

CyberGym leans defensive, and on the offence-heavy ExploitBench GLM-5.3 still sits behind the closed frontier. So this is not "the best hacker model ever shipped." What it is — a top-of-the-open-field coding model whose single biggest generational jump is in finding and exploiting software flaws, about to be released under a licence that lets anyone run it locally with no rate limit and no usage policy — is precisely the combination that makes a two-week hold on the weights look less like caution theatre and more like the correct call.

Why the delay matters more than the benchmark

Open-weights releases have mostly been a story about access: cheaper inference, no vendor lock-in, models you can run on your own hardware. Holding the weights back like this is a rare move for a leading open-weights lab, and a clear first for Z.ai — which shipped GLM-5.2's weights to Hugging Face within days of launch. It is not unprecedented: OpenAI famously staged the GPT-2 release back in 2019 over misuse worries. But for the current crop of open labs, most of them racing weights out to stay competitive, a self-imposed safety hold on your strongest model is the exception, not the rule.

That makes it a small precedent with a large shadow. Once the weights are public, they are public forever — there is no recall, no patch, no rate limit you can throttle after the fact. A staged release buys a safety review; it does not buy a second chance. Watching whether the rest of the open field — Qwen, DeepSeek, Meta — adopts the same restraint, or races the weights out to stay ahead, will tell you more about where open AI is heading than any single benchmark on the chart.

For now, GLM-5.3 is available through Z.ai's GLM Coding Plan and its ZCode IDE. The weights follow when the safety evaluation is done. If you build on open models, that sequence — capability first, weights second, safety review in between — is the one worth getting used to.

Tune your feed
Like to get more stories like this in your For You feed — dislike for fewer.
Sources
Des Okoro — Research Correspondent. Des covers the research desk — papers, benchmarks, and breakthroughs — and translates how the tech really works under the hood. Spot something wrong? Tell me and I'll correct it in public.
Got a question about this?

Ask Relay — he reads every question himself and replies personally by email.

Ask Relay →