99MODELS

کیمی K2.6

Previous Kimi generation; strong agentic tool use at low cost.

تاریخ اجراء: 20 اپریل، 2026

384.00

فی 10 لاکھ آؤٹ پٹ Tokens

ان پٹ: ₹91.20 فی 10 لاکھ Tokens

فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔

تکنیکی تفصیلات

Context ونڈو
2,62,144 Tokens
زیادہ سے زیادہ آؤٹ پٹ
16,384 Tokens
سپورٹ
ٹیکسٹ, تصاویر
Reasoning
پہلے سے آن
ٹول کا استعمال
ہاں
سٹرکچرڈ آؤٹ پٹ
ہاں
کوڈ ایگزیکیوشن
نہیں
ذہانت کا درجہ
54 میں سے #32
ویلیو کا درجہ
54 میں سے #26

قیمتیں

قیمتیں
فی 10 lakh TokensINRUSD
ان پٹ91.20$0.95
آؤٹ پٹ384.00$4.00
کیشڈ ان پٹ33.60$0.35

Benchmarks

تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔

  • 91.1%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 37.5%

    HLE

    Humanity's Last Exam

  • 76.0%

    IFBench

    IFBench - precise instruction following

  • 81.0%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 43.9%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 65.9%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 76.7%

    SWE-bench Verified

    SWE-bench Verified - real-world bug fixing (Epoch AI run)

  • 22.1%

    OSWorld 2

    OSWorld 2 - agentic computer use (partial credit)

  • 1509

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

کیمی K2.6 کے بارے میں

Moonshot AI's open-source multimodal agentic model, advancing long-horizon coding, coding-driven design, proactive autonomous execution and swarm-based task orchestration. It is a 1T-parameter Mixture-of-Experts with 32B activated parameters and a vision encoder, supporting both visual and text input across a 256K context. Thinking is the default mode and can be switched off for an instant-response mode.

K2.6 is the release where Moonshot's argument moves from single tasks to sessions that run for most of a working day, and two of its own examples set the scale. Asked to download and run a small language model locally on a Mac, it implemented and optimised inference in Zig, a deliberately niche choice, across more than 4,000 tool calls, over twelve hours of continuous execution and fourteen iterations, taking throughput from about 15 to about 193 tokens per second - which the lab measures as roughly 20% faster than a widely used desktop runtime. In a second run it spent thirteen hours on an eight-year-old open-source financial matching engine, worked through twelve optimisation strategies and more than a thousand tool calls, read CPU and allocation flame graphs to find the real bottleneck, reconfigured the core thread topology, and reports throughput gains of 185% and 133% on the engine's two headline measures.

The other new capability is orchestration. Moonshot says the agent swarm now scales to 300 sub-agents across 4,000 coordinated steps, against 100 and 1,500 in the K2.5 research preview, and that it can turn a supplied document, spreadsheet or deck into a reusable skill that keeps the original's structure and style. Its own reliability team ran a K2.6-backed agent autonomously for five days handling monitoring, incident response and system operations, which is the kind of claim that is easy to state and hard to fake.

On benchmarks Moonshot reports 80.2 on SWE-bench Verified, 76.7 on the multilingual set, 58.6 on SWE-bench Pro, 66.7 on Terminal-Bench 2.0 and 89.6 on LiveCodeBench v6, with 90.5 on GPQA-Diamond and 92.5 F1 on DeepSearchQA. Coding-driven design is the other emphasis: complete front-end interfaces from a single prompt with deliberate layout, interaction and scroll-triggered animation, and simple full-stack flows spanning authentication, user interaction and database work.

The same table is where the limits sit, and they are worth reading. Moonshot's own numbers put K2.6 behind GPT-5.4 and Claude Opus 4.6 on several agentic sets, at 27.9 against 33.3 and 33.0 on APEX-Agents, 50.0 against 54.6 on Toolathlon and 55.9 against 62.5 on MCPMark, and at 34.7 on Humanity's Last Exam without tools it is behind every competing model in its own table. Broad reasoning is not what this release was aimed at.

One caveat is unusual enough to be worth repeating: because the weights are open, how the model is served changes what you get. Moonshot notes that reproducing its published numbers requires the official API and points at its own vendor-verification project for judging third-party hosts. That is the honest footnote attached to every open-weight release - the weights are the same everywhere, the serving is not. The copy served here is Moonshot's.

لانچ کے وقت Moonshot AI کا بیان

Long-horizon coding
Moonshot reports twelve-hour and thirteen-hour unattended runs on real codebases, taking a local inference implementation from about 15 to about 193 tokens per second across 4,000 tool calls.
Coding-driven design
Complete front-end interfaces from a single prompt with deliberate layout, interaction and animation, extending to simple full-stack flows covering authentication, user interaction and database work.
Agent swarms, scaled up
The swarm now runs 300 sub-agents across 4,000 coordinated steps, against 100 and 1,500 in the K2.5 preview, and can turn supplied documents into reusable skills that keep their structure and style.
Agents that run unattended
Moonshot's own reliability team ran a K2.6-backed agent autonomously for five days on monitoring, incident response and operations, from alert through to resolution.
What it is not for
The lab reports it behind GPT-5.4 and Claude Opus 4.6 on several agentic evaluations, and behind every competing model in its own table on Humanity's Last Exam without tools. Broad reasoning was not the target.
Serving affects the scores
Moonshot states that reproducing its published results needs the official API and publishes a vendor-verification project for judging third-party hosts of the open weights.

ہندوستانی زبانیں

کیمی K2.6 6 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے زبان کا انتخاب کریں۔

Frequently Asked Questions

کیمی K2.6 کے بارے میں اکثر پوچھے جانے والے سوالات۔

کیمی K2.6 کب جاری ہوا تھا؟

Moonshot AI نے کیمی K2.6 کو 20 اپریل، 2026 کو جاری کیا۔

کیمی K2.6 کس نے بنایا ہے؟

کیمی K2.6 کو Moonshot AI نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔

کیمی K2.6 کتنا ذہین ہے؟

یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 32 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔

کیمی K2.6 کا کتنا خرچ آتا ہے؟

استعمال کی لاگت ₹91.20 فی 10 لاکھ ان پٹ Tokens اور ₹384.00 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔

امریکی ڈالر میں کیمی K2.6 کی قیمت کیا ہے؟

فراہم کنندہ $0.95 فی 10 لاکھ ان پٹ Tokens اور $4.00 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔

کیمی K2.6 کتنی لمبی گفتگو یاد رکھ سکتا ہے؟

اس کی Context حد 2.6 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔

کیا کیمی K2.6 مناسب قیمت میں بہترین کارکردگی دیتا ہے؟

یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 26 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔

کیا کیمی K2.6 جواب دینے سے پہلے سوچتا ہے؟

پہلے سے آن۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔

کیمی K2.6 کن ہندوستانی زبانوں میں جواب دیتا ہے؟

یہ 6 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے اپنی زبان منتخب کریں۔

Moonshot AI کے مزید ماڈلز

کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026