Kimi K2.6
Previous Kimi generation; strong agentic tool use at low cost.
Released Apr 20, 2026
₹384.00
per 10 lakh output tokens
Input: ₹91.20 per 10 lakh tokens
Billed at provider rates converted at ₹96 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
91.1%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
37.5%
HLE
Humanity's Last Exam
76.0%
IFBench
IFBench - precise instruction following
81.0%
Long Context
Long Context Reasoning - reasoning over long inputs
43.9%
Terminal-Bench Hard
Terminal-Bench Hard - agentic terminal tasks
65.9%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
76.7%
SWE-bench Verified
SWE-bench Verified - real-world bug fixing (Epoch AI run)
22.1%
OSWorld 2
OSWorld 2 - agentic computer use (partial credit)
1509
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
About Kimi K2.6
Moonshot AI's open-source multimodal agentic model, advancing long-horizon coding, coding-driven design, proactive autonomous execution and swarm-based task orchestration. It is a 1T-parameter Mixture-of-Experts with 32B activated parameters and a vision encoder, supporting both visual and text input across a 256K context. Thinking is the default mode and can be switched off for an instant-response mode.
K2.6 is the release where Moonshot's argument moves from single tasks to sessions that run for most of a working day, and two of its own examples set the scale. Asked to download and run a small language model locally on a Mac, it implemented and optimised inference in Zig, a deliberately niche choice, across more than 4,000 tool calls, over twelve hours of continuous execution and fourteen iterations, taking throughput from about 15 to about 193 tokens per second - which the lab measures as roughly 20% faster than a widely used desktop runtime. In a second run it spent thirteen hours on an eight-year-old open-source financial matching engine, worked through twelve optimisation strategies and more than a thousand tool calls, read CPU and allocation flame graphs to find the real bottleneck, reconfigured the core thread topology, and reports throughput gains of 185% and 133% on the engine's two headline measures.
The other new capability is orchestration. Moonshot says the agent swarm now scales to 300 sub-agents across 4,000 coordinated steps, against 100 and 1,500 in the K2.5 research preview, and that it can turn a supplied document, spreadsheet or deck into a reusable skill that keeps the original's structure and style. Its own reliability team ran a K2.6-backed agent autonomously for five days handling monitoring, incident response and system operations, which is the kind of claim that is easy to state and hard to fake.
On benchmarks Moonshot reports 80.2 on SWE-bench Verified, 76.7 on the multilingual set, 58.6 on SWE-bench Pro, 66.7 on Terminal-Bench 2.0 and 89.6 on LiveCodeBench v6, with 90.5 on GPQA-Diamond and 92.5 F1 on DeepSearchQA. Coding-driven design is the other emphasis: complete front-end interfaces from a single prompt with deliberate layout, interaction and scroll-triggered animation, and simple full-stack flows spanning authentication, user interaction and database work.
The same table is where the limits sit, and they are worth reading. Moonshot's own numbers put K2.6 behind GPT-5.4 and Claude Opus 4.6 on several agentic sets, at 27.9 against 33.3 and 33.0 on APEX-Agents, 50.0 against 54.6 on Toolathlon and 55.9 against 62.5 on MCPMark, and at 34.7 on Humanity's Last Exam without tools it is behind every competing model in its own table. Broad reasoning is not what this release was aimed at.
One caveat is unusual enough to be worth repeating: because the weights are open, how the model is served changes what you get. Moonshot notes that reproducing its published numbers requires the official API and points at its own vendor-verification project for judging third-party hosts. That is the honest footnote attached to every open-weight release - the weights are the same everywhere, the serving is not. The copy served here is Moonshot's.
What Moonshot AI announced at launch
- Long-horizon coding
- Moonshot reports twelve-hour and thirteen-hour unattended runs on real codebases, taking a local inference implementation from about 15 to about 193 tokens per second across 4,000 tool calls.
- Coding-driven design
- Complete front-end interfaces from a single prompt with deliberate layout, interaction and animation, extending to simple full-stack flows covering authentication, user interaction and database work.
- Agent swarms, scaled up
- The swarm now runs 300 sub-agents across 4,000 coordinated steps, against 100 and 1,500 in the K2.5 preview, and can turn supplied documents into reusable skills that keep their structure and style.
- Agents that run unattended
- Moonshot's own reliability team ran a K2.6-backed agent autonomously for five days on monitoring, incident response and operations, from alert through to resolution.
- What it is not for
- The lab reports it behind GPT-5.4 and Claude Opus 4.6 on several agentic evaluations, and behind every competing model in its own table on Humanity's Last Exam without tools. Broad reasoning was not the target.
- Serving affects the scores
- Moonshot states that reproducing its published results needs the official API and publishes a vendor-verification project for judging third-party hosts of the open weights.
Indian languages
Kimi K2.6 answers in 6 Indian languages. Choose a language from the menu beside the message box to receive replies in it.
Frequently Asked Questions
Frequently asked questions about Kimi K2.6.
When was Kimi K2.6 released?
Moonshot AI released Kimi K2.6 on Apr 20, 2026.
Who built Kimi K2.6?
Kimi K2.6 is developed by Moonshot AI. 99Models AI connects directly to it at the provider's published rate.
How intelligent is Kimi K2.6?
It is ranked 32 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does Kimi K2.6 cost?
Usage costs ₹91.20 per million input tokens and ₹384.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is Kimi K2.6 pricing in US dollars?
The provider charges $0.95 per million input tokens and $4.00 per million output tokens. Rupee rates are converted at ₹96 per US dollar.
How long a conversation can Kimi K2.6 hold?
Its context window is 2.6 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does Kimi K2.6 rank for value?
It ranks 26 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does Kimi K2.6 support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
Which Indian languages does Kimi K2.6 support?
It answers in 6 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.
More models from Moonshot AI
Catalog updated Sep 9, 2026