GLM 5.1
Previous GLM-5 generation; solid general reasoning.
റിലീസ് തീയതി: 2026 ഏപ്രി 7
₹464.64
10 ലക്ഷം ഔട്ട്പുട്ട് Tokens-ന്
ഇൻപുട്ട്: 10 ലക്ഷം Tokens-ന് ₹147.84
0% മാർക്ക്അപ്പിൽ, ഒരു ഡോളറിന് ₹96 എന്ന നിരക്കിലാണ് ഈടാക്കുന്നത്.
നിരക്കുകൾ
ബെഞ്ച്മാർക്കുകൾ
റേറ്റിംഗ് എന്ന് രേഖപ്പെടുത്താത്ത സ്കോറുകൾ ശതമാനത്തിലാണ്. എല്ലാ ബെഞ്ച്മാർക്കുകളും സ്വതന്ത്രമായി പരിശോധിച്ചവയാണ്.
86.8%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
30.1%
HLE
Humanity's Last Exam
76.3%
IFBench
IFBench - precise instruction following
73.7%
Long Context
Long Context Reasoning - reasoning over long inputs
43.2%
Terminal-Bench Hard
Terminal-Bench Hard - agentic terminal tasks
61.8%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
74.2%
SWE-bench Verified
SWE-bench Verified - real-world bug fixing (Epoch AI run)
1509
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
GLM 5.1-നെ കുറിച്ച്
Z.ai's previous flagship for agentic engineering, with significantly stronger coding than its predecessor. It is designed for long-horizon autonomous work, sustaining optimisation over hundreds of rounds and thousands of tool calls rather than plateauing early, with performance that improves as the task lengthens -- Z.ai cites sustained execution for up to eight hours. Thinking is on by default and can be disabled.
The framing of the 5.1 post is a claim about time rather than about peak scores. Z.ai argues that earlier models, its own GLM-5 included, exhaust their repertoire early: they apply the familiar techniques, take a quick gain, then plateau, and giving them longer does not help. GLM-5.1 was built so that extra runtime keeps paying, and the post spends most of its length demonstrating that rather than listing results.
It uses three tasks with progressively less structured feedback. The first is an open-source vector-database challenge scored purely on queries per second at 95% recall, normally run inside a 50-turn tool budget where the best result on record was about 3,500 queries per second. Z.ai restructured it as an outer optimisation loop and reports GLM-5.1 still finding real improvements past 600 iterations and 6,000 tool calls, finishing at 21.5 thousand queries per second. The trajectory is a staircase: six structural rewrites, each chosen by the model after reading its own benchmark logs, with recall temporarily broken around each transition and then restored.
The second is more honest about the ceiling. On the hardest level of a GPU kernel benchmark, where whole architectures have to be optimised end to end, Z.ai reports GLM-5.1 reaching a 3.6 times geometric-mean speedup and still improving late in the run, while naming Claude Opus 4.6 as the stronger model in that setting at 4.2 times with headroom left. The third has no metric at all: an eight-hour run building a Linux-style desktop as a web application, wrapped in a harness that asks the model after each round what is still missing.
The benchmark numbers behind that: Z.ai reports 58.4 on SWE-bench Pro, 42.7 on natural-language-to-repository generation and 69.0 on Terminal-Bench 2.0 with its best harness, against GLM-5's 55.1, 35.9 and 56.2, plus 68.7 on the CyberGym vulnerability-reproduction set against GLM-5's 48.3.
The post ends on what is not solved, which is the part worth reading before you plan a long autonomous run on it: escaping a local optimum earlier once incremental tuning stops paying, holding coherence across execution traces of thousands of tool calls, and, the one Z.ai flags as most important, reliable self-evaluation where there is no number to optimise against. It calls this a first step in that direction. The weights are MIT-licensed and can be self-hosted; the copy served here is Z.ai's.
ലോഞ്ചിംഗ് വേളയിൽ Z.ai വ്യക്തമാക്കിയത്
- Complex software engineering
- Z.ai reports state-of-the-art results on SWE-bench Pro at 58.4, ahead of its own GLM-5 at 55.1 and of the closed models it compares against on that particular set.
- Real-world terminal tasks
- On Terminal-Bench 2.0 the lab reports 69.0 with its best harness against 56.2 for GLM-5, one of the widest generational gaps in the post.
- Repository generation
- On long-horizon repository generation from a natural-language brief, Z.ai reports 42.7 against 35.9 for GLM-5, with Claude Opus 4.6 still ahead at 49.8.
- Cybersecurity reproduction
- On CyberGym, which asks a model to reproduce known software vulnerabilities, the lab reports 68.7 against 48.3 for GLM-5 and ahead of the closed models in its table.
- Runtime that keeps paying
- The central claim: on a vector-search optimisation task the lab reports the model still improving past 600 iterations and 6,000 tool calls, ending roughly six times better than the best single-session result.
- What it is not for
- Z.ai names three unsolved problems: escaping local optima earlier, holding coherence over thousands of tool calls, and self-evaluation on tasks with no numeric target. It is behind Claude Opus 4.6 on kernel optimisation.
ഇന്ത്യൻ ഭാഷകൾ
GLM 5.1 11 ഇന്ത്യൻ ഭാഷകളിൽ മറുപടി നൽകും. മെസ്സേജ് ബോക്സിന് അടുത്തുള്ള മെനുവിൽ നിന്ന് ആവശ്യമുള്ള ഭാഷ തിരഞ്ഞെടുക്കാം.
Frequently Asked Questions
GLM 5.1 സംബന്ധിച്ച പ്രധാന ചോദ്യങ്ങളും ഉത്തരങ്ങളും.
GLM 5.1 എപ്പോഴാണ് റിലീസ് ചെയ്തത്?
Z.ai 2026 ഏപ്രി 7-ൽ GLM 5.1 പുറത്തിറക്കി.
GLM 5.1 വികസിപ്പിച്ചത് ആരാണ്?
GLM 5.1 വികസിപ്പിച്ചത് Z.ai ആണ്. 99Models AI ഒറിജിനൽ പ്രൊവൈഡർ നിരക്കിൽ തന്നെ ഇതിലേക്ക് കണക്ട് ചെയ്യുന്നു.
GLM 5.1 എത്രത്തോളം കാര്യക്ഷമമാണ്?
ഇന്റലിജൻസ് റാങ്കിംഗിൽ 54 മോഡലുകളിൽ 38-ാം സ്ഥാനത്താണ് ഇത്. ബെഞ്ച്മാർക്ക് സ്കോറുകൾ മുകളിലുള്ള പാനലിൽ കാണാം.
GLM 5.1 ഉപയോഗിക്കാൻ എത്ര ചെലവാകും?
10 ലക്ഷം ഇൻപുട്ട് Tokens-ന് ₹147.84 രൂപയും ഔട്ട്പുട്ടിന് ₹464.64 രൂപയുമാണ് അധിക മാർക്ക്അപ്പില്ലാത്ത നിരക്ക്. സബ്സ്ക്രിപ്ഷനില്ല, ഉപയോഗിക്കുന്നതിന് മാത്രം പണമടയ്ക്കുക.
GLM 5.1 API-യുടെ ഡോളർ നിരക്ക് എത്രയാണ്?
10 ലക്ഷം ഇൻപുട്ട് Tokens-ന് $1.54 ഡോളറും 10 ലക്ഷം ഔട്ട്പുട്ട് Tokens-ന് $4.84 ഡോളറുമാണ് നിരക്ക്. ഈ പേജിലെ രൂപ നിരക്കുകൾ ഡോളറിന് ₹96 എന്ന നിരക്കിൽ മാറ്റിയതാണ്.
GLM 5.1-ന് എത്ര നീളമുള്ള സംഭാഷണം ഓർത്തുനിൽക്കാനാകും?
ഇതിന്റെ Context window 2 lakh Tokens ആണ്. ഒരുമിച്ച് നൽകുന്ന സംഭാഷണങ്ങളും ഫയലുകളും ഉൾപ്പെടെ ഇതിൽ വായിക്കാൻ സാധിക്കും.
GLM 5.1 മികച്ച വാല്യൂ നൽകുന്ന ഒന്നാണോ?
പെർഫോമൻസും നിരക്കും അടിസ്ഥാനമാക്കിയുള്ള വാല്യൂ റാങ്കിംഗിൽ 54 മോഡലുകളിൽ 42-ാം സ്ഥാനത്താണ് ഇത്.
മറുപടി നൽകുന്നതിന് മുൻപ് GLM 5.1 Reasoning നടത്തുമോ?
ഡിഫോൾട്ടായി ഓൺ ആണ്. Reasoning പിന്തുണയ്ക്കുന്ന മോഡലുകളിൽ ചിന്തിക്കുന്നതിന്റെ വ്യാപ്തി മെസ്സേജ് ബോക്സിൽ ക്രമീകരിക്കാം.
GLM 5.1 ഏതൊക്കെ ഇന്ത്യൻ ഭാഷകളിൽ മറുപടി നൽകും?
11 ഇന്ത്യൻ ഭാഷകളിൽ ഇത് മറുപടി നൽകും. മെസ്സേജ് ബോക്സിന് അടുത്തുള്ള മെനുവിൽ നിന്ന് ഭാഷ തിരഞ്ഞെടുക്കാം.
Z.ai-ൽ നിന്നുള്ള മറ്റ് Models
കാറ്റലോഗ് പുതുക്കിയത്: 2026 സെപ്റ്റം 9