GLM 5.3 فلیش
نیاCheapest GLM-5.3; native vision and a million-token context.
تاریخ اجراء: 26 اگست، 2026
₹48.00
فی 10 لاکھ آؤٹ پٹ Tokens
ان پٹ: ₹14.40 فی 10 لاکھ Tokens
فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔
تکنیکی تفصیلات
- Context ونڈو
- 13,10,720 Tokens
- زیادہ سے زیادہ آؤٹ پٹ
- 1,31,072 Tokens
- سپورٹ
- ٹیکسٹ, تصاویر, ویڈیو
- Reasoning
- پہلے سے آن
- ایفرٹ لیولز
- low, high, max
- ٹول کا استعمال
- ہاں
- سٹرکچرڈ آؤٹ پٹ
- ہاں
- کوڈ ایگزیکیوشن
- نہیں
- ذہانت کا درجہ
- 54 میں سے #17
- ویلیو کا درجہ
- 54 میں سے #1
قیمتیں
Benchmarks
تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔
91.2%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
39.9%
HLE
Humanity's Last Exam
51.6%
SciCode
SciCode - scientific code generation
80.0%
Long Context
Long Context Reasoning - reasoning over long inputs
84.3%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
1604
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
GLM 5.3 فلیش کے بارے میں
Z.ai's efficiency-tier 5.3, and the only model in the generation with a direct US-hosted route. It is natively multimodal over text, images and video, and pairs a hybrid sparse-and-linear attention architecture with a million-token window, which is what lets it hold accurate long-context behaviour while cutting the compute a dense attention stack would spend at that length. Z.ai positions it for efficient coding and long-horizon agent work rather than frontier reasoning, and it prices roughly an order of magnitude below the full 5.3. Thinking is mandatory and always on, with three effort levels.
Z.ai did not distil this one down from the larger 5.3. It starts from a newly trained base model whose architecture and training recipe were redesigned around getting more capability out of less compute, and the shape of that redesign is easy to read. Against GLM-4.5, at a similar total size of 320B versus 355B, it roughly halves both the active parameters per token (18B against 32B) and the layer count (45 against 92). Alongside the hybrid attention it adds IndexPool, which compresses four indexer key vectors into one so the indexer's latency and memory cost stay bounded at a million tokens, and Manifold-Constrained Hyper-Connections to improve scaling efficiency. Z.ai measures three times less attention compute and a 4.4 times smaller cache than the full GLM-5.3.
The headline claim is a price claim. Z.ai reports GLM-5.3-Flash ahead of GLM-5.2 across benchmarks and real workloads at a tenth of the cost, by 63.4 to 46.2 on DeepSWE v1.1 and 48.8 to 26.2 on Zapier's AutomationBench, and on the lab's own Z.ai Code Bench at max effort it lands within half a point of Claude Opus 4.8 (29.0 against 29.5). Before the release the model spent a week running anonymously under a codename in public coding tools so the lab could collect feedback on it without a name attached.
Vision is the part Z.ai argues is structural rather than decorative. Its case is that for frontend work, game development, 3D simulation and slide generation the deliverable is a rendered thing, and many failures only surface once it is rendered, so the model has to be able to decide when to look. It was trained on trajectories that require it to inspect its own output and revise it, with reinforcement learning against environment feedback for frontend coding and agent-based verification grounded in real user flows. The lab extends the same loop to documents, spreadsheets, dashboards and presentations, where the argument is that the model reads the artefacts of a task instead of asking you to describe them.
The other unusual thing about this release is where it ran. Z.ai served the launch week entirely on a large cluster of Chinese AI chips, with a dedicated inference engine built on SGLang, aggressive memory and quantisation work, and an encode-prefill-decode disaggregated architecture that schedules the three stages independently. It reports a threefold end-to-end serving improvement over its own baseline on that hardware, reaching per-token cost comparable to mainstream GPUs. The weights are published on Hugging Face and run under SGLang, vLLM and TokenSpeed, so this is a model you could host yourself; the copy served here is Z.ai's.
Z.ai is clear that this is the cost-performance model rather than the peak one. It notes that the cache is still slightly larger than two comparable recent open models and calls that room for improvement, and it closes by saying the recipe is now being scaled to larger models, with what the team learned building this one already shaping the next frontier release.
لانچ کے وقت Z.ai کا بیان
- Flash cost, higher scores
- Z.ai reports GLM-5.3-Flash ahead of GLM-5.2 across six coding and agentic benchmarks at a tenth of the price, by 63.4 to 46.2 on DeepSWE v1.1 and 48.8 to 26.2 on AutomationBench.
- Architecture for efficiency
- At a similar total size to GLM-4.5 it halves both the active parameters and the layer count, and the lab measures three times less attention compute and a 4.4 times smaller cache than the full GLM-5.3.
- Vision in the coding loop
- Trained to render its own output, inspect what a user would actually see and revise the artefact, which Z.ai argues is the only way to catch frontend, game and 3D failures that never appear in the code itself.
- Knowledge work, not just code
- The same visual reasoning is pointed at documents, spreadsheets, presentations, dashboards and interfaces, so the model interprets a task from its artefacts rather than from a written description of them.
- Served on domestic silicon
- Z.ai ran the entire launch week on a cluster of Chinese AI chips, reporting a threefold serving improvement over its own baseline on that hardware and per-token cost comparable to mainstream GPUs.
- What it is not for
- This is the cost-performance model, not the frontier one. Z.ai notes its cache is still larger than two comparable open models and says the recipe is now being scaled up to a larger frontier release.
ہندوستانی زبانیں
GLM 5.3 فلیش 12 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے زبان کا انتخاب کریں۔
Frequently Asked Questions
GLM 5.3 فلیش کے بارے میں اکثر پوچھے جانے والے سوالات۔
GLM 5.3 فلیش کب جاری ہوا تھا؟
Z.ai نے GLM 5.3 فلیش کو 26 اگست، 2026 کو جاری کیا۔
GLM 5.3 فلیش کس نے بنایا ہے؟
GLM 5.3 فلیش کو Z.ai نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔
GLM 5.3 فلیش کتنا ذہین ہے؟
یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 17 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔
GLM 5.3 فلیش کا کتنا خرچ آتا ہے؟
استعمال کی لاگت ₹14.40 فی 10 لاکھ ان پٹ Tokens اور ₹48.00 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔
امریکی ڈالر میں GLM 5.3 فلیش کی قیمت کیا ہے؟
فراہم کنندہ $0.15 فی 10 لاکھ ان پٹ Tokens اور $0.50 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔
GLM 5.3 فلیش کتنی لمبی گفتگو یاد رکھ سکتا ہے؟
اس کی Context حد 13.1 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔
کیا GLM 5.3 فلیش مناسب قیمت میں بہترین کارکردگی دیتا ہے؟
یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 1 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔
کیا GLM 5.3 فلیش جواب دینے سے پہلے سوچتا ہے؟
پہلے سے آن۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔
GLM 5.3 فلیش کن ہندوستانی زبانوں میں جواب دیتا ہے؟
یہ 12 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے اپنی زبان منتخب کریں۔
Z.ai کے مزید ماڈلز
کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026