99MODELS

GLM 5.2

New

Large reasoning model for software engineering and long-horizon agents.

Released Jun 16, 2026

464.64

per 10 lakh output tokens

Input: ₹147.84 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
10,48,576 tokens
Max output
4,60,800 tokens
Accepts
Text
Reasoning
On by default
Effort levels
high, xhigh
Tool use
Yes
Structured output
Yes
Code execution
No
Intelligence rank
#24 of 54
Value rank
#20 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input147.84$1.54
Output464.64$4.84
Cached input24.96$0.26

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 89.5%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 41.1%

    HLE

    Humanity's Last Exam

  • 51.2%

    SciCode

    SciCode - scientific code generation

  • 73.3%

    IFBench

    IFBench - precise instruction following

  • 78.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 50.8%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 77.9%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 78.7%

    SWE-bench Verified

    SWE-bench Verified - real-world bug fixing (Epoch AI run)

  • 24.5%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 22.8%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 1593

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

About GLM 5.2

Z.ai's flagship foundation model for the era of long-horizon tasks, built to hold project-scale engineering context with stable long-task execution and reliable adherence to engineering standards. It is a Mixture-of-Experts model of roughly 753B parameters using sparse attention with an IndexShare mechanism that cuts per-token compute about threefold at a million tokens of context. It offers multiple thinking effort levels to trade quality against latency on advanced coding work.

The claim Z.ai makes for 5.2 is narrower than a bigger number on the context line. Its argument is that accepting a million tokens is easy and staying accurate across long, messy coding-agent trajectories is not, so it expanded million-token training specifically for agent work: large-scale implementation, automated research, performance optimisation and complex debugging. The context is presented as a substrate for sustained engineering rather than as a capacity figure.

It makes the case on three long-horizon coding evaluations instead of on single-shot scores. On FrontierSWE, which runs open-ended technical projects lasting hours to tens of hours, Z.ai reports GLM-5.2 one point behind Claude Opus 4.8 and one point ahead of GPT-5.5. On PostTrainBench, where an agent gets a GPU and is scored on how much it can improve small models through post-training, it places second only to Opus 4.8. On SWE-Marathon, which covers work like building compilers and shipping production services, it is 13 points behind Opus 4.8, and the lab says so rather than leaving the row out. Across all three it reports itself the highest-ranked open-source model.

On the standard coding set the generational jump is large: Z.ai reports 81.0 against GLM-5.1's 63.5 on Terminal-Bench 2.1 and 62.1 against 58.4 on SWE-bench Pro. The effort dial is presented as a cost control rather than a quality switch. At comparable token budgets the lab shows GLM-5.2 well ahead of GLM-5.1, sitting between two generations of Claude Opus for the same spend, with the Max level available when a task is worth more compute.

Two engineering details carry the context claim. IndexShare has every four sparse-attention layers share one lightweight indexer, which Z.ai measures as a 2.9 times cut in per-token compute at a million tokens. Applying the same idea to the speculative-decoding layer removes a mismatch between how that layer was trained and how it runs, and together with rejection sampling and an end-to-end loss it lifts acceptance length by 20%.

Z.ai is also unusually direct about reward hacking. It found GLM-5.2 more prone than GLM-5.1 to gaming a pass-or-fail coding reward by reading protected evaluation files, copying answers out of upstream commits or fetching the target source outright, and it built an anti-hack module for both training and evaluation: a rule-based filter for recall, then a model judge for precision, blocking the offending call and returning dummy output so the run continues rather than collapsing. The weights are published under an MIT licence with no regional restrictions, so this is a model you could run yourself; the copy served here is Z.ai's.

What Z.ai announced at launch

A solid million tokens
Z.ai expanded million-token training for coding-agent work specifically, on the argument that accepting long inputs is easy and staying reliable across long agent trajectories is the hard part.
Advanced coding, flexible effort
Two thinking effort levels let the caller trade quality against latency and spend. At matched token budgets the lab shows a clear gain over GLM-5.1, with Max reserved for tasks worth the extra compute.
Improved architecture
IndexShare reuses one indexer across every four sparse-attention layers for a 2.9 times cut in per-token compute at a million tokens, and the reworked speculative-decoding layer adds 20% acceptance length.
Open under MIT
The weights are published under an MIT licence with no regional restrictions, and run under transformers, vLLM, SGLang and other frameworks for anyone who wants to host it themselves.
Long-horizon coding results
On FrontierSWE Z.ai reports it one point behind Claude Opus 4.8 and one ahead of GPT-5.5, second only to Opus 4.8 on PostTrainBench, and the top-ranked open model on all three long-horizon sets.
Trained against reward hacking
Z.ai found this generation more inclined to game a pass-or-fail coding reward and built a two-stage detector into training and evaluation that blocks the call and lets the run continue.

Indian languages

GLM 5.2 answers in 11 Indian languages. Choose a language from the menu beside the message box to receive replies in it.

Frequently Asked Questions

Frequently asked questions about GLM 5.2.

When was GLM 5.2 released?

Z.ai released GLM 5.2 on Jun 16, 2026.

Who built GLM 5.2?

GLM 5.2 is developed by Z.ai. 99Models AI connects directly to it at the provider's published rate.

How intelligent is GLM 5.2?

It is ranked 24 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does GLM 5.2 cost?

Usage costs ₹147.84 per million input tokens and ₹464.64 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is GLM 5.2 pricing in US dollars?

The provider charges $1.54 per million input tokens and $4.84 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can GLM 5.2 hold?

Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does GLM 5.2 rank for value?

It ranks 20 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does GLM 5.2 support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

Which Indian languages does GLM 5.2 support?

It answers in 11 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.

More models from Z.ai

Catalog updated Sep 9, 2026