GLM 5.3
NewFlagship GLM generation for long-horizon engineering and agents.
Released Aug 18, 2026
₹633.60
per 10 lakh output tokens
Input: ₹201.60 per 10 lakh tokens
Billed at provider rates converted at ₹96 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
91.7%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
42.3%
HLE
Humanity's Last Exam
59.0%
SciCode
SciCode - scientific code generation
79.7%
Long Context
Long Context Reasoning - reasoning over long inputs
83.9%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
1599
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
About GLM 5.3
Z.ai's flagship generation for coding and long-horizon agentic work, and an unusual release in that it shares GLM-5.2's base model entirely -- every gain comes from scaled-up post-training. Z.ai reports it as the strongest open-weights coding model it has measured, with roughly a 50 percent improvement over 5.2 on its own coding evaluation, and it reaches those scores while spending fewer output tokens per task. A capability the lab says grew faster than expected is vulnerability analysis: it reasons across multiple stages of exploitation rather than spotting isolated flaws. Thinking is mandatory, at three effort levels.
The unusual thing about GLM-5.3 is what did not change. It runs on GLM-5.2's base model, unmodified - Z.ai's own summary of the release is that scaling post-training is all it did. Everything below is therefore a claim about how much is still available after pretraining stops, which makes this release readable as an experiment as much as a product.
The coding gains are the headline. Z.ai calls it the strongest open-weights coding model it has measured, and the numbers it reports are large enough to be worth stating individually: Terminal-Bench 3.0 from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents' Last Exam from 23.8 to 28.5. It attributes the durability of those gains to reinforcement-learning strategies carried over from 5.2, including compaction, which it says is what makes the improvements hold on long-horizon tasks rather than only on short ones.
The efficiency dimension is where Z.ai draws its sharpest comparison, because it measures score against tokens spent rather than score alone. On its in-house Z.ai Code Bench - a private benchmark it built to reduce contamination from public test sets - it reports 34.5% at roughly 75K output tokens per task at Max effort, against 23.4% at 96K for GLM-5.2: better and cheaper at once. At High effort it reports 31.4% at around 50K output tokens, ahead of Claude Opus 4.8's 29.5% at 120K. Z.ai is also explicit about where it stops: it says GLM-5.3 remains behind Claude Fable 5, which reaches 39.5% at Max effort.
The part Z.ai says surprised it is security. It added vulnerability-discovery data and environments to the post-training mix expecting the model to get better at spotting flaws, and reports that the capability kept developing faster than anticipated as training scaled - the model began reasoning across multiple stages of exploitation and forming coherent plans for complete chains, rather than identifying isolated bugs. Run against real codebases with security teams, and after expert review and deduplication, it reports 2,436 vulnerabilities found across 269 projects, 1,097 of them medium-to-high severity, spanning kernels, operating systems, browser engines, infrastructure, web applications and network protocols. Many had gone unnoticed for years; the oldest dated back around four decades.
Thinking is mandatory on this model and runs at three effort levels. Z.ai says the weights would follow the API release by about two weeks, after safety evaluation and hardening.
What Z.ai announced at launch
- Post-training alone
- GLM-5.3 shares GLM-5.2's base model unchanged. Z.ai's own framing is that scaling post-training is all it did, which makes the gains below a measurement of how much is left after pretraining ends.
- Frontier open-weights coding
- Z.ai reports Terminal-Bench 3.0 rising from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9 and Agents' Last Exam from 23.8 to 28.5, and calls it the strongest open-weights coding model it has measured.
- Better and cheaper per task
- On its private Z.ai Code Bench the lab reports 34.5% at about 75K output tokens at Max effort against 23.4% at 96K for GLM-5.2, and at High effort 31.4% at around 50K tokens against Claude Opus 4.8 at 29.5% with 120K.
- Cyber capability that outgrew training
- Z.ai says vulnerability analysis developed faster than it expected as training scaled: the model reasons across multiple stages of exploitation and forms complete chains rather than spotting isolated flaws.
- Findings on real code
- Run against real codebases with security teams, and after expert review and deduplication, Z.ai reports 2,436 vulnerabilities across 269 projects, 1,097 of them medium-to-high severity, some decades old.
- What it is not for
- Z.ai names its own ceiling: on its coding benchmark GLM-5.3 remains behind Claude Fable 5, which reaches 39.5% at Max effort. Thinking is mandatory here, so there is no cheap non-reasoning mode to fall back to.
Frequently Asked Questions
Frequently asked questions about GLM 5.3.
When was GLM 5.3 released?
Z.ai released GLM 5.3 on Aug 18, 2026.
Who built GLM 5.3?
GLM 5.3 is developed by Z.ai. 99Models AI connects directly to it at the provider's published rate.
How intelligent is GLM 5.3?
It is ranked 10 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does GLM 5.3 cost?
Usage costs ₹201.60 per million input tokens and ₹633.60 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is GLM 5.3 pricing in US dollars?
The provider charges $2.10 per million input tokens and $6.60 per million output tokens. Rupee rates are converted at ₹96 per US dollar.
How long a conversation can GLM 5.3 hold?
Its context window is 13.1 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does GLM 5.3 rank for value?
It ranks 13 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does GLM 5.3 support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
More models from Z.ai
Catalog updated Sep 9, 2026