GLM 5.3 Flash
NewCheapest GLM-5.3; native vision and a million-token context.
Released Aug 26, 2026
₹48.00
per 10 lakh output tokens
Input: ₹14.40 per 10 lakh tokens
Billed at provider rates converted at ₹96 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
91.2%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
39.9%
HLE
Humanity's Last Exam
51.6%
SciCode
SciCode - scientific code generation
80.0%
Long Context
Long Context Reasoning - reasoning over long inputs
84.3%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
1604
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
About GLM 5.3 Flash
Z.ai's efficiency-tier 5.3, and the only model in the generation with a direct US-hosted route. It is natively multimodal over text, images and video, and pairs a hybrid sparse-and-linear attention architecture with a million-token window, which is what lets it hold accurate long-context behaviour while cutting the compute a dense attention stack would spend at that length. Z.ai positions it for efficient coding and long-horizon agent work rather than frontier reasoning, and it prices roughly an order of magnitude below the full 5.3. Thinking is mandatory and always on, with three effort levels.
Z.ai did not distil this one down from the larger 5.3. It starts from a newly trained base model whose architecture and training recipe were redesigned around getting more capability out of less compute, and the shape of that redesign is easy to read. Against GLM-4.5, at a similar total size of 320B versus 355B, it roughly halves both the active parameters per token (18B against 32B) and the layer count (45 against 92). Alongside the hybrid attention it adds IndexPool, which compresses four indexer key vectors into one so the indexer's latency and memory cost stay bounded at a million tokens, and Manifold-Constrained Hyper-Connections to improve scaling efficiency. Z.ai measures three times less attention compute and a 4.4 times smaller cache than the full GLM-5.3.
The headline claim is a price claim. Z.ai reports GLM-5.3-Flash ahead of GLM-5.2 across benchmarks and real workloads at a tenth of the cost, by 63.4 to 46.2 on DeepSWE v1.1 and 48.8 to 26.2 on Zapier's AutomationBench, and on the lab's own Z.ai Code Bench at max effort it lands within half a point of Claude Opus 4.8 (29.0 against 29.5). Before the release the model spent a week running anonymously under a codename in public coding tools so the lab could collect feedback on it without a name attached.
Vision is the part Z.ai argues is structural rather than decorative. Its case is that for frontend work, game development, 3D simulation and slide generation the deliverable is a rendered thing, and many failures only surface once it is rendered, so the model has to be able to decide when to look. It was trained on trajectories that require it to inspect its own output and revise it, with reinforcement learning against environment feedback for frontend coding and agent-based verification grounded in real user flows. The lab extends the same loop to documents, spreadsheets, dashboards and presentations, where the argument is that the model reads the artefacts of a task instead of asking you to describe them.
The other unusual thing about this release is where it ran. Z.ai served the launch week entirely on a large cluster of Chinese AI chips, with a dedicated inference engine built on SGLang, aggressive memory and quantisation work, and an encode-prefill-decode disaggregated architecture that schedules the three stages independently. It reports a threefold end-to-end serving improvement over its own baseline on that hardware, reaching per-token cost comparable to mainstream GPUs. The weights are published on Hugging Face and run under SGLang, vLLM and TokenSpeed, so this is a model you could host yourself; the copy served here is Z.ai's.
Z.ai is clear that this is the cost-performance model rather than the peak one. It notes that the cache is still slightly larger than two comparable recent open models and calls that room for improvement, and it closes by saying the recipe is now being scaled to larger models, with what the team learned building this one already shaping the next frontier release.
What Z.ai announced at launch
- Flash cost, higher scores
- Z.ai reports GLM-5.3-Flash ahead of GLM-5.2 across six coding and agentic benchmarks at a tenth of the price, by 63.4 to 46.2 on DeepSWE v1.1 and 48.8 to 26.2 on AutomationBench.
- Architecture for efficiency
- At a similar total size to GLM-4.5 it halves both the active parameters and the layer count, and the lab measures three times less attention compute and a 4.4 times smaller cache than the full GLM-5.3.
- Vision in the coding loop
- Trained to render its own output, inspect what a user would actually see and revise the artefact, which Z.ai argues is the only way to catch frontend, game and 3D failures that never appear in the code itself.
- Knowledge work, not just code
- The same visual reasoning is pointed at documents, spreadsheets, presentations, dashboards and interfaces, so the model interprets a task from its artefacts rather than from a written description of them.
- Served on domestic silicon
- Z.ai ran the entire launch week on a cluster of Chinese AI chips, reporting a threefold serving improvement over its own baseline on that hardware and per-token cost comparable to mainstream GPUs.
- What it is not for
- This is the cost-performance model, not the frontier one. Z.ai notes its cache is still larger than two comparable open models and says the recipe is now being scaled up to a larger frontier release.
Indian languages
GLM 5.3 Flash answers in 12 Indian languages. Choose a language from the menu beside the message box to receive replies in it.
Frequently Asked Questions
Frequently asked questions about GLM 5.3 Flash.
When was GLM 5.3 Flash released?
Z.ai released GLM 5.3 Flash on Aug 26, 2026.
Who built GLM 5.3 Flash?
GLM 5.3 Flash is developed by Z.ai. 99Models AI connects directly to it at the provider's published rate.
How intelligent is GLM 5.3 Flash?
It is ranked 17 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does GLM 5.3 Flash cost?
Usage costs ₹14.40 per million input tokens and ₹48.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is GLM 5.3 Flash pricing in US dollars?
The provider charges $0.15 per million input tokens and $0.50 per million output tokens. Rupee rates are converted at ₹96 per US dollar.
How long a conversation can GLM 5.3 Flash hold?
Its context window is 13.1 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does GLM 5.3 Flash rank for value?
It ranks 1 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does GLM 5.3 Flash support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
Which Indian languages does GLM 5.3 Flash support?
It answers in 12 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.
More models from Z.ai
Catalog updated Sep 9, 2026