99MODELS

کیمی K3

نیا

2.8T open-weight multimodal reasoner for long-horizon agentic work.

تاریخ اجراء: 16 جولائی، 2026

1,584.00

فی 10 لاکھ آؤٹ پٹ Tokens

ان پٹ: ₹316.80 فی 10 لاکھ Tokens

فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔

تکنیکی تفصیلات

Context ونڈو
10,48,576 Tokens
زیادہ سے زیادہ آؤٹ پٹ
9,43,718 Tokens
سپورٹ
ٹیکسٹ, تصاویر, ویڈیو
Reasoning
پہلے سے آن
ایفرٹ لیولز
low, high, max
ٹول کا استعمال
ہاں
سٹرکچرڈ آؤٹ پٹ
ہاں
کوڈ ایگزیکیوشن
نہیں
پیرامیٹرز
2.8T MoE
ذہانت کا درجہ
54 میں سے #9
ویلیو کا درجہ
54 میں سے #21

قیمتیں

قیمتیں
فی 10 lakh TokensINRUSD
ان پٹ316.80$3.30
آؤٹ پٹ1,584.00$16.50
کیشڈ ان پٹ31.68$0.33

Benchmarks

تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔

  • 93.5%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 46.9%

    HLE

    Humanity's Last Exam

  • 59.5%

    SciCode

    SciCode - scientific code generation

  • 88.7%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 85.0%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 44.2%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 60.4%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 1674

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

کیمی K3 کے بارے میں

Moonshot AI's most capable model: an open-weight, natively multimodal agentic model with 2.8T total parameters and roughly 104B activated across 896 experts. It is built on Kimi Delta Attention and Attention Residuals with a Stable LatentMoE framework, giving roughly 2.5 times better scaling efficiency than Kimi K2, and pairs native visual understanding with a million-token context. It is designed for frontier work -- software engineering, knowledge work and deep reasoning -- and always reasons.

Moonshot's own framing is that the parameter count is not the point. The architecture changes are what the lab argues for: Kimi Delta Attention as an efficient base for scaling attention, and Attention Residuals, which retrieve representations selectively across depth instead of accumulating them uniformly. Sparsity was pushed further too, with a Stable LatentMoE framework effectively activating 16 of 896 experts, and the routing work that makes that stable at scale - allocation derived directly from router-score quantiles rather than from a hand-tuned balancing term, and an optimiser that treats each attention head independently. The claim Moonshot puts on all of it together is roughly 2.5 times better scaling efficiency than Kimi K2: compute converted into capability, rather than parameters converted into a headline.

The case for the model is made mostly in case studies. Given 24 hours per task in identical sandboxes, Moonshot reports K3 competitive with Claude Fable 5 and substantially ahead of Claude Opus 4.8, GPT 5.6 Sol and GPT 5.5 at optimising GPU kernels across two hardware families, and notes that late in K3's own development an early version handled most of the team's kernel work. Asked to build a GPU programming system from scratch it produced a compact Triton-like compiler with its own tile-level intermediate representation, optimisation passes and a code-generation path down to PTX, which the lab says matches or beats the established stack on supported benchmarks and trains a small transformer end to end. In a single 48-hour autonomous run it designed, optimised and verified a chip for a nano model built on its own architecture, closing timing at 100 MHz inside four square millimetres.

The published table backs that with 88.3 on Terminal-Bench 2.1, 42.0 on SWE Marathon and 77.8 on Program Bench, the highest in its compared set on the last two, alongside 93.5 on GPQA-Diamond and 91.2 on BrowseComp. Multimodality is native rather than bolted on, and Moonshot leans on it for video in particular: it points to K3 cutting its own launch teaser from 56 source clips with beat-synchronised edits, and producing an animated explainer of its own architecture.

The limitations section is the part to read before wiring it into anything. K3 was trained in preserved-thinking mode, so a harness that fails to pass the full thinking history back, or a session handed over to K3 from another model mid-way, can make generation highly unstable. It is also, in the lab's own words, prone to excessive proactiveness: trained hard on long-horizon work, it tends to decide on your behalf when the intent is ambiguous, and needs explicit behavioural constraints in the system prompt if it must stay inside defined boundaries. Moonshot states plainly that overall performance still trails Claude Fable 5 and GPT 5.6 Sol, with a noticeable gap in user experience.

The weights are open. Moonshot committed to publishing the full set by 27 July 2026 and contributed a prefix-caching implementation for its new attention design upstream so the model can be served efficiently outside the lab's own stack. Self-hosting it is a serious undertaking rather than a download, though: the recommendation is deployment on supernode configurations of 64 or more accelerators. The copy served here is Moonshot's.

لانچ کے وقت Moonshot AI کا بیان

An open 3T-class model
Moonshot presents K3 as the first open model at 2.8 trillion parameters, and claims roughly 2.5 times better scaling efficiency than Kimi K2 from its new attention and mixture-of-experts design.
Kernels, compilers, silicon
The lab reports K3 competitive with Claude Fable 5 at GPU kernel optimisation, building a working Triton-like compiler from scratch, and designing and verifying a small chip in one 48-hour run.
Long-horizon coding
On its published table Moonshot reports 88.3 on Terminal-Bench 2.1, plus the top scores in its compared set on SWE Marathon at 42.0 and Program Bench at 77.8.
Knowledge work end to end
Native vision is used for research reports with generated charts and interactive narratives, for editing video, and for producing infographic-style presentations rather than only reading images.
Preserved thinking is required
K3 was trained with its full thinking history carried forward. A harness that drops it, or a session switched to K3 mid-way from another model, can make output highly unstable.
What it is not for
Moonshot warns the model improvises on your behalf when intent is ambiguous and needs explicit constraints, and says overall performance and user experience still trail Claude Fable 5 and GPT 5.6 Sol.

ہندوستانی زبانیں

کیمی K3 7 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے زبان کا انتخاب کریں۔

Frequently Asked Questions

کیمی K3 کے بارے میں اکثر پوچھے جانے والے سوالات۔

کیمی K3 کب جاری ہوا تھا؟

Moonshot AI نے کیمی K3 کو 16 جولائی، 2026 کو جاری کیا۔

کیمی K3 کس نے بنایا ہے؟

کیمی K3 کو Moonshot AI نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔

کیمی K3 کتنا ذہین ہے؟

یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 9 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔

کیمی K3 کا کتنا خرچ آتا ہے؟

استعمال کی لاگت ₹316.80 فی 10 لاکھ ان پٹ Tokens اور ₹1,584.00 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔

امریکی ڈالر میں کیمی K3 کی قیمت کیا ہے؟

فراہم کنندہ $3.30 فی 10 لاکھ ان پٹ Tokens اور $16.50 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔

کیمی K3 کتنی لمبی گفتگو یاد رکھ سکتا ہے؟

اس کی Context حد 10.5 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔

کیا کیمی K3 مناسب قیمت میں بہترین کارکردگی دیتا ہے؟

یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 21 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔

کیا کیمی K3 جواب دینے سے پہلے سوچتا ہے؟

پہلے سے آن۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔

کیمی K3 کن ہندوستانی زبانوں میں جواب دیتا ہے؟

یہ 7 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے اپنی زبان منتخب کریں۔

Moonshot AI کے مزید ماڈلز

کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026