کیمی K3
نیا2.8T open-weight multimodal reasoner for long-horizon agentic work.
تاریخ اجراء: 16 جولائی، 2026
₹1,584.00
فی 10 لاکھ آؤٹ پٹ Tokens
ان پٹ: ₹316.80 فی 10 لاکھ Tokens
فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔
تکنیکی تفصیلات
- Context ونڈو
- 10,48,576 Tokens
- زیادہ سے زیادہ آؤٹ پٹ
- 9,43,718 Tokens
- سپورٹ
- ٹیکسٹ, تصاویر, ویڈیو
- Reasoning
- پہلے سے آن
- ایفرٹ لیولز
- low, high, max
- ٹول کا استعمال
- ہاں
- سٹرکچرڈ آؤٹ پٹ
- ہاں
- کوڈ ایگزیکیوشن
- نہیں
- پیرامیٹرز
- 2.8T MoE
- ذہانت کا درجہ
- 54 میں سے #9
- ویلیو کا درجہ
- 54 میں سے #21
قیمتیں
Benchmarks
تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔
93.5%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
46.9%
HLE
Humanity's Last Exam
59.5%
SciCode
SciCode - scientific code generation
88.7%
Long Context
Long Context Reasoning - reasoning over long inputs
85.0%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
44.2%
FrontierCode
FrontierCode - long-horizon production coding tasks
60.4%
ARC-AGI-2
ARC-AGI-2 - abstract reasoning on novel puzzles
1674
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
کیمی K3 کے بارے میں
Moonshot AI's most capable model: an open-weight, natively multimodal agentic model with 2.8T total parameters and roughly 104B activated across 896 experts. It is built on Kimi Delta Attention and Attention Residuals with a Stable LatentMoE framework, giving roughly 2.5 times better scaling efficiency than Kimi K2, and pairs native visual understanding with a million-token context. It is designed for frontier work -- software engineering, knowledge work and deep reasoning -- and always reasons.
Moonshot's own framing is that the parameter count is not the point. The architecture changes are what the lab argues for: Kimi Delta Attention as an efficient base for scaling attention, and Attention Residuals, which retrieve representations selectively across depth instead of accumulating them uniformly. Sparsity was pushed further too, with a Stable LatentMoE framework effectively activating 16 of 896 experts, and the routing work that makes that stable at scale - allocation derived directly from router-score quantiles rather than from a hand-tuned balancing term, and an optimiser that treats each attention head independently. The claim Moonshot puts on all of it together is roughly 2.5 times better scaling efficiency than Kimi K2: compute converted into capability, rather than parameters converted into a headline.
The case for the model is made mostly in case studies. Given 24 hours per task in identical sandboxes, Moonshot reports K3 competitive with Claude Fable 5 and substantially ahead of Claude Opus 4.8, GPT 5.6 Sol and GPT 5.5 at optimising GPU kernels across two hardware families, and notes that late in K3's own development an early version handled most of the team's kernel work. Asked to build a GPU programming system from scratch it produced a compact Triton-like compiler with its own tile-level intermediate representation, optimisation passes and a code-generation path down to PTX, which the lab says matches or beats the established stack on supported benchmarks and trains a small transformer end to end. In a single 48-hour autonomous run it designed, optimised and verified a chip for a nano model built on its own architecture, closing timing at 100 MHz inside four square millimetres.
The published table backs that with 88.3 on Terminal-Bench 2.1, 42.0 on SWE Marathon and 77.8 on Program Bench, the highest in its compared set on the last two, alongside 93.5 on GPQA-Diamond and 91.2 on BrowseComp. Multimodality is native rather than bolted on, and Moonshot leans on it for video in particular: it points to K3 cutting its own launch teaser from 56 source clips with beat-synchronised edits, and producing an animated explainer of its own architecture.
The limitations section is the part to read before wiring it into anything. K3 was trained in preserved-thinking mode, so a harness that fails to pass the full thinking history back, or a session handed over to K3 from another model mid-way, can make generation highly unstable. It is also, in the lab's own words, prone to excessive proactiveness: trained hard on long-horizon work, it tends to decide on your behalf when the intent is ambiguous, and needs explicit behavioural constraints in the system prompt if it must stay inside defined boundaries. Moonshot states plainly that overall performance still trails Claude Fable 5 and GPT 5.6 Sol, with a noticeable gap in user experience.
The weights are open. Moonshot committed to publishing the full set by 27 July 2026 and contributed a prefix-caching implementation for its new attention design upstream so the model can be served efficiently outside the lab's own stack. Self-hosting it is a serious undertaking rather than a download, though: the recommendation is deployment on supernode configurations of 64 or more accelerators. The copy served here is Moonshot's.
لانچ کے وقت Moonshot AI کا بیان
- An open 3T-class model
- Moonshot presents K3 as the first open model at 2.8 trillion parameters, and claims roughly 2.5 times better scaling efficiency than Kimi K2 from its new attention and mixture-of-experts design.
- Kernels, compilers, silicon
- The lab reports K3 competitive with Claude Fable 5 at GPU kernel optimisation, building a working Triton-like compiler from scratch, and designing and verifying a small chip in one 48-hour run.
- Long-horizon coding
- On its published table Moonshot reports 88.3 on Terminal-Bench 2.1, plus the top scores in its compared set on SWE Marathon at 42.0 and Program Bench at 77.8.
- Knowledge work end to end
- Native vision is used for research reports with generated charts and interactive narratives, for editing video, and for producing infographic-style presentations rather than only reading images.
- Preserved thinking is required
- K3 was trained with its full thinking history carried forward. A harness that drops it, or a session switched to K3 mid-way from another model, can make output highly unstable.
- What it is not for
- Moonshot warns the model improvises on your behalf when intent is ambiguous and needs explicit constraints, and says overall performance and user experience still trail Claude Fable 5 and GPT 5.6 Sol.
ہندوستانی زبانیں
کیمی K3 7 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے زبان کا انتخاب کریں۔
Frequently Asked Questions
کیمی K3 کے بارے میں اکثر پوچھے جانے والے سوالات۔
کیمی K3 کب جاری ہوا تھا؟
Moonshot AI نے کیمی K3 کو 16 جولائی، 2026 کو جاری کیا۔
کیمی K3 کس نے بنایا ہے؟
کیمی K3 کو Moonshot AI نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔
کیمی K3 کتنا ذہین ہے؟
یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 9 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔
کیمی K3 کا کتنا خرچ آتا ہے؟
استعمال کی لاگت ₹316.80 فی 10 لاکھ ان پٹ Tokens اور ₹1,584.00 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔
امریکی ڈالر میں کیمی K3 کی قیمت کیا ہے؟
فراہم کنندہ $3.30 فی 10 لاکھ ان پٹ Tokens اور $16.50 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔
کیمی K3 کتنی لمبی گفتگو یاد رکھ سکتا ہے؟
اس کی Context حد 10.5 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔
کیا کیمی K3 مناسب قیمت میں بہترین کارکردگی دیتا ہے؟
یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 21 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔
کیا کیمی K3 جواب دینے سے پہلے سوچتا ہے؟
پہلے سے آن۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔
کیمی K3 کن ہندوستانی زبانوں میں جواب دیتا ہے؟
یہ 7 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے اپنی زبان منتخب کریں۔
Moonshot AI کے مزید ماڈلز
کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026