Kimi K3
New2.8T open-weight multimodal reasoner for long-horizon agentic work.
Released Jul 16, 2026
₹1,584.00
per 10 lakh output tokens
Input: ₹316.80 per 10 lakh tokens
Billed at provider rates converted at ₹96 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
93.5%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
46.9%
HLE
Humanity's Last Exam
59.5%
SciCode
SciCode - scientific code generation
88.7%
Long Context
Long Context Reasoning - reasoning over long inputs
85.0%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
44.2%
FrontierCode
FrontierCode - long-horizon production coding tasks
60.4%
ARC-AGI-2
ARC-AGI-2 - abstract reasoning on novel puzzles
1674
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
About Kimi K3
Moonshot AI's most capable model: an open-weight, natively multimodal agentic model with 2.8T total parameters and roughly 104B activated across 896 experts. It is built on Kimi Delta Attention and Attention Residuals with a Stable LatentMoE framework, giving roughly 2.5 times better scaling efficiency than Kimi K2, and pairs native visual understanding with a million-token context. It is designed for frontier work -- software engineering, knowledge work and deep reasoning -- and always reasons.
Moonshot's own framing is that the parameter count is not the point. The architecture changes are what the lab argues for: Kimi Delta Attention as an efficient base for scaling attention, and Attention Residuals, which retrieve representations selectively across depth instead of accumulating them uniformly. Sparsity was pushed further too, with a Stable LatentMoE framework effectively activating 16 of 896 experts, and the routing work that makes that stable at scale - allocation derived directly from router-score quantiles rather than from a hand-tuned balancing term, and an optimiser that treats each attention head independently. The claim Moonshot puts on all of it together is roughly 2.5 times better scaling efficiency than Kimi K2: compute converted into capability, rather than parameters converted into a headline.
The case for the model is made mostly in case studies. Given 24 hours per task in identical sandboxes, Moonshot reports K3 competitive with Claude Fable 5 and substantially ahead of Claude Opus 4.8, GPT 5.6 Sol and GPT 5.5 at optimising GPU kernels across two hardware families, and notes that late in K3's own development an early version handled most of the team's kernel work. Asked to build a GPU programming system from scratch it produced a compact Triton-like compiler with its own tile-level intermediate representation, optimisation passes and a code-generation path down to PTX, which the lab says matches or beats the established stack on supported benchmarks and trains a small transformer end to end. In a single 48-hour autonomous run it designed, optimised and verified a chip for a nano model built on its own architecture, closing timing at 100 MHz inside four square millimetres.
The published table backs that with 88.3 on Terminal-Bench 2.1, 42.0 on SWE Marathon and 77.8 on Program Bench, the highest in its compared set on the last two, alongside 93.5 on GPQA-Diamond and 91.2 on BrowseComp. Multimodality is native rather than bolted on, and Moonshot leans on it for video in particular: it points to K3 cutting its own launch teaser from 56 source clips with beat-synchronised edits, and producing an animated explainer of its own architecture.
The limitations section is the part to read before wiring it into anything. K3 was trained in preserved-thinking mode, so a harness that fails to pass the full thinking history back, or a session handed over to K3 from another model mid-way, can make generation highly unstable. It is also, in the lab's own words, prone to excessive proactiveness: trained hard on long-horizon work, it tends to decide on your behalf when the intent is ambiguous, and needs explicit behavioural constraints in the system prompt if it must stay inside defined boundaries. Moonshot states plainly that overall performance still trails Claude Fable 5 and GPT 5.6 Sol, with a noticeable gap in user experience.
The weights are open. Moonshot committed to publishing the full set by 27 July 2026 and contributed a prefix-caching implementation for its new attention design upstream so the model can be served efficiently outside the lab's own stack. Self-hosting it is a serious undertaking rather than a download, though: the recommendation is deployment on supernode configurations of 64 or more accelerators. The copy served here is Moonshot's.
What Moonshot AI announced at launch
- An open 3T-class model
- Moonshot presents K3 as the first open model at 2.8 trillion parameters, and claims roughly 2.5 times better scaling efficiency than Kimi K2 from its new attention and mixture-of-experts design.
- Kernels, compilers, silicon
- The lab reports K3 competitive with Claude Fable 5 at GPU kernel optimisation, building a working Triton-like compiler from scratch, and designing and verifying a small chip in one 48-hour run.
- Long-horizon coding
- On its published table Moonshot reports 88.3 on Terminal-Bench 2.1, plus the top scores in its compared set on SWE Marathon at 42.0 and Program Bench at 77.8.
- Knowledge work end to end
- Native vision is used for research reports with generated charts and interactive narratives, for editing video, and for producing infographic-style presentations rather than only reading images.
- Preserved thinking is required
- K3 was trained with its full thinking history carried forward. A harness that drops it, or a session switched to K3 mid-way from another model, can make output highly unstable.
- What it is not for
- Moonshot warns the model improvises on your behalf when intent is ambiguous and needs explicit constraints, and says overall performance and user experience still trail Claude Fable 5 and GPT 5.6 Sol.
Indian languages
Kimi K3 answers in 7 Indian languages. Choose a language from the menu beside the message box to receive replies in it.
Frequently Asked Questions
Frequently asked questions about Kimi K3.
When was Kimi K3 released?
Moonshot AI released Kimi K3 on Jul 16, 2026.
Who built Kimi K3?
Kimi K3 is developed by Moonshot AI. 99Models AI connects directly to it at the provider's published rate.
How intelligent is Kimi K3?
It is ranked 9 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does Kimi K3 cost?
Usage costs ₹316.80 per million input tokens and ₹1,584.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is Kimi K3 pricing in US dollars?
The provider charges $3.30 per million input tokens and $16.50 per million output tokens. Rupee rates are converted at ₹96 per US dollar.
How long a conversation can Kimi K3 hold?
Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does Kimi K3 rank for value?
It ranks 21 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does Kimi K3 support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
Which Indian languages does Kimi K3 support?
It answers in 7 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.
More models from Moonshot AI
Catalog updated Sep 9, 2026