99MODELS

GLM 5.1

Previous GLM-5 generation; solid general reasoning.

ରିଲିଜ୍ ତାରିଖ ଅପ୍ରେଲ 7, 2026

464.64

ପ୍ରତି 10 ଲକ୍ଷ output Tokens

Input: ପ୍ରତି 10 ଲକ୍ଷ Tokens ପାଇଁ ₹147.84

0% ମାର୍କଅପ୍ ସହିତ ₹96 ପ୍ରତି US ଡଲାର ହିସାବରେ ପ୍ରୋଭାଇଡର୍ ରେଟ୍ ରେ ବିଲ୍ କରାଯାଇଛି।

ସ୍ପେସିଫିକେସନ୍

Context window
2,04,800 Tokens
ସର୍ବାଧିକ Output
65,536 Tokens
ଗ୍ରହଣ କରେ
ଟେକ୍ସଟ୍
Reasoning
ଡିଫଲ୍ଟ ଭାବରେ On
Tool ବ୍ୟବହାର
ହଁ
Structured output
ହଁ
Code execution
ନାହିଁ
Intelligence rank
54 ମଧ୍ୟରୁ #38

ମୂଲ୍ୟ

ମୂଲ୍ୟ
ପ୍ରତି 10 ଲକ୍ଷ TokensINRUSD
Input147.84$1.54
Output464.64$4.84
Cached input27.46$0.29

Benchmarks

ରେଟିଂ ଭାବେ ଚିହ୍ନିତ ନ ହେଲେ ସ୍କୋରଗୁଡ଼ିକ ପ୍ରତିଶତ ଅଟେ। ସମସ୍ତ benchmark ସ୍ୱତନ୍ତ୍ର ଭାବରେ ମପାଯାଇଛି।

  • 86.8%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 30.1%

    HLE

    Humanity's Last Exam

  • 76.3%

    IFBench

    IFBench - precise instruction following

  • 73.7%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 43.2%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 61.8%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 74.2%

    SWE-bench Verified

    SWE-bench Verified - real-world bug fixing (Epoch AI run)

  • 1509

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

GLM 5.1 ବିଷୟରେ

Z.ai's previous flagship for agentic engineering, with significantly stronger coding than its predecessor. It is designed for long-horizon autonomous work, sustaining optimisation over hundreds of rounds and thousands of tool calls rather than plateauing early, with performance that improves as the task lengthens -- Z.ai cites sustained execution for up to eight hours. Thinking is on by default and can be disabled.

The framing of the 5.1 post is a claim about time rather than about peak scores. Z.ai argues that earlier models, its own GLM-5 included, exhaust their repertoire early: they apply the familiar techniques, take a quick gain, then plateau, and giving them longer does not help. GLM-5.1 was built so that extra runtime keeps paying, and the post spends most of its length demonstrating that rather than listing results.

It uses three tasks with progressively less structured feedback. The first is an open-source vector-database challenge scored purely on queries per second at 95% recall, normally run inside a 50-turn tool budget where the best result on record was about 3,500 queries per second. Z.ai restructured it as an outer optimisation loop and reports GLM-5.1 still finding real improvements past 600 iterations and 6,000 tool calls, finishing at 21.5 thousand queries per second. The trajectory is a staircase: six structural rewrites, each chosen by the model after reading its own benchmark logs, with recall temporarily broken around each transition and then restored.

The second is more honest about the ceiling. On the hardest level of a GPU kernel benchmark, where whole architectures have to be optimised end to end, Z.ai reports GLM-5.1 reaching a 3.6 times geometric-mean speedup and still improving late in the run, while naming Claude Opus 4.6 as the stronger model in that setting at 4.2 times with headroom left. The third has no metric at all: an eight-hour run building a Linux-style desktop as a web application, wrapped in a harness that asks the model after each round what is still missing.

The benchmark numbers behind that: Z.ai reports 58.4 on SWE-bench Pro, 42.7 on natural-language-to-repository generation and 69.0 on Terminal-Bench 2.0 with its best harness, against GLM-5's 55.1, 35.9 and 56.2, plus 68.7 on the CyberGym vulnerability-reproduction set against GLM-5's 48.3.

The post ends on what is not solved, which is the part worth reading before you plan a long autonomous run on it: escaping a local optimum earlier once incremental tuning stops paying, holding coherence across execution traces of thousands of tool calls, and, the one Z.ai flags as most important, reliable self-evaluation where there is no number to optimise against. It calls this a first step in that direction. The weights are MIT-licensed and can be self-hosted; the copy served here is Z.ai's.

ଲଞ୍ଚ ସମୟରେ Z.ai ଯାହା କହିଥିଲା

Complex software engineering
Z.ai reports state-of-the-art results on SWE-bench Pro at 58.4, ahead of its own GLM-5 at 55.1 and of the closed models it compares against on that particular set.
Real-world terminal tasks
On Terminal-Bench 2.0 the lab reports 69.0 with its best harness against 56.2 for GLM-5, one of the widest generational gaps in the post.
Repository generation
On long-horizon repository generation from a natural-language brief, Z.ai reports 42.7 against 35.9 for GLM-5, with Claude Opus 4.6 still ahead at 49.8.
Cybersecurity reproduction
On CyberGym, which asks a model to reproduce known software vulnerabilities, the lab reports 68.7 against 48.3 for GLM-5 and ahead of the closed models in its table.
Runtime that keeps paying
The central claim: on a vector-search optimisation task the lab reports the model still improving past 600 iterations and 6,000 tool calls, ending roughly six times better than the best single-session result.
What it is not for
Z.ai names three unsolved problems: escaping local optima earlier, holding coherence over thousands of tool calls, and self-evaluation on tasks with no numeric target. It is behind Claude Opus 4.6 on kernel optimisation.

ଭାରତୀୟ ଭାଷା

GLM 5.1 11 ଟି ଭାରତୀୟ ଭାଷାରେ ଉତ୍ତର ଦିଏ। ସେହି ଭାଷାରେ ଉତ୍ତର ପାଇବା ପାଇଁ ମେସେଜ୍ ବକ୍ସ ପାଖରେ ଥିବା ମେନୁରୁ ଭାଷା ବାଛନ୍ତୁ।

Frequently Asked Questions

GLM 5.1 ବିଷୟରେ ବାରମ୍ବାର ପଚରାଯାଉଥିବା ପ୍ରଶ୍ନ।

GLM 5.1 କେବେ ଲଞ୍ଚ ହୋଇଥିଲା?

Z.ai ଅପ୍ରେଲ 7, 2026 ରେ GLM 5.1 ଲଞ୍ଚ କରିଥିଲା।

GLM 5.1 କିଏ ତିଆରି କରିଛି?

GLM 5.1 କୁ Z.ai ତିଆରି କରିଛି। 99Models ଏହା ସହ ସିଧାସଳଖ ପ୍ରୋଭାଇଡରଙ୍କ ନିର୍ଦ୍ଧାରିତ ରେଟ୍ ରେ ସଂଯୋଗ କରେ।

GLM 5.1 କେତେ ଶକ୍ତିଶାଳୀ?

ଇଣ୍ଟେଲିଜେନ୍ସ ତାଲିକାରେ 54 ଟି chat Model ମଧ୍ୟରୁ ଏହାର ରାଙ୍କ୍ 38। ସ୍ୱାଧୀନ ବେଞ୍ଚମାର୍କ ସ୍କୋର ଆଧାରରେ ଏହା ସ୍ଥିର କରାଯାଇଛି, ଯାହା ଉପରେ ଥିବା ଟେବୁଲରେ ଉପଲବ୍ଧ।

GLM 5.1 ର ମୂଲ୍ୟ କେତେ?

ବ୍ୟବହାର ଖର୍ଚ୍ଚ ପ୍ରତି 10 ଲକ୍ଷ ଇନପୁଟ୍ Tokens ପାଇଁ ₹147.84 ଏବଂ ଆଉଟପୁଟ୍ Tokens ପାଇଁ ₹464.64, 0% ମାର୍କଅପ୍ ସହିତ। କୌଣସି ସବସ୍କ୍ରିପସନ୍ ନାହିଁ; ଆପଣ ଯେତିକି ବ୍ୟବହାର କରିବେ ସେତିକି ପେମେଣ୍ଟ କରିବେ।

US ଡଲାରରେ GLM 5.1 ର ମୂଲ୍ୟ କେତେ?

ପ୍ରୋଭାଇଡର୍ ପ୍ରତି 10 ଲକ୍ଷ ଇନପୁଟ୍ Tokens ପାଇଁ $1.54 ଏବଂ ଆଉଟପୁଟ୍ Tokens ପାଇଁ $4.84 ଚାର୍ଜ କରେ। ଏହି ପୃଷ୍ଠାର ଟଙ୍କା ମୂଲ୍ୟ ₹96 ପ୍ରତି ଡଲାର ହିସାବରେ ରୂପାନ୍ତରିତ।

GLM 5.1 କେତେ ଲମ୍ବା କଥାବାର୍ତ୍ତା ମନେ ରଖିପାରିବ?

ଏହାର Context window ହେଉଛି 2 lakh Tokens। ଗୋଟିଏ request ରେ ଏହା ସମୁଦାୟ କଥାବାର୍ତ୍ତା ଏବଂ ସଂଲଗ୍ନ ଫାଇଲ୍ ପଢ଼ିପାରିବ।

ମୂଲ୍ୟ ହିସାବରେ GLM 5.1 କେତେ ଭଲ?

ଭ୍ୟାଲୁ ରାଙ୍କିଙ୍ଗରେ 54 ଟି Model ମଧ୍ୟରୁ ଏହାର ସ୍ଥାନ 42। ଏହି ରାଙ୍କିଙ୍ଗ ବେଞ୍ଚମାର୍କ କ୍ଷମତା ଏବଂ Token ଖର୍ଚ୍ଚକୁ ତୁଳନା କରି ସ୍ଥିର କରାଯାଏ।

GLM 5.1 କ’ଣ Reasoning ସପୋର୍ଟ କରେ?

ଡିଫଲ୍ଟ ଭାବରେ On। ଯେଉଁଠାରେ Reasoning ଉପଲବ୍ଧ, ଆପଣ ମେସେଜ୍ ବକ୍ସରେ thinking effort ସ୍ତର ସେଟ୍ କରିପାରିବେ।

GLM 5.1 କେଉଁ ଭାରତୀୟ ଭାଷାରେ ଉତ୍ତର ଦେଇପାରେ?

ଏହା 11 ଟି ଭାରତୀୟ ଭାଷାରେ ଉତ୍ତର ଦିଏ। ମେସେଜ୍ ବକ୍ସ ପାଖରେ ଥିବା ମେନୁରୁ ଆପଣଙ୍କ ପସନ୍ଦର ଭାଷା ବାଛନ୍ତୁ।

Z.ai ରୁ ଅନ୍ୟାନ୍ୟ Models

କାଟାଲଗ୍ ଅପଡେଟ୍ ହୋଇଛି: ସେପ୍ଟେମ୍ବର 9, 2026