କିମି K3
ନୂଆ2.8T open-weight multimodal reasoner for long-horizon agentic work.
ରିଲିଜ୍ ତାରିଖ ଜୁଲାଇ 16, 2026
₹1,584.00
ପ୍ରତି 10 ଲକ୍ଷ output Tokens
Input: ପ୍ରତି 10 ଲକ୍ଷ Tokens ପାଇଁ ₹316.80
0% ମାର୍କଅପ୍ ସହିତ ₹96 ପ୍ରତି US ଡଲାର ହିସାବରେ ପ୍ରୋଭାଇଡର୍ ରେଟ୍ ରେ ବିଲ୍ କରାଯାଇଛି।
ସ୍ପେସିଫିକେସନ୍
- Context window
- 10,48,576 Tokens
- ସର୍ବାଧିକ Output
- 9,43,718 Tokens
- ଗ୍ରହଣ କରେ
- ଟେକ୍ସଟ୍, ଛବି, Video
- Reasoning
- ଡିଫଲ୍ଟ ଭାବରେ On
- Effort ସ୍ତର
- low, high, max
- Tool ବ୍ୟବହାର
- ହଁ
- Structured output
- ହଁ
- Code execution
- ନାହିଁ
- Parameters
- 2.8T MoE
- Intelligence rank
- 54 ମଧ୍ୟରୁ #9
- Value rank
- 54 ମଧ୍ୟରୁ #21
ମୂଲ୍ୟ
Benchmarks
ରେଟିଂ ଭାବେ ଚିହ୍ନିତ ନ ହେଲେ ସ୍କୋରଗୁଡ଼ିକ ପ୍ରତିଶତ ଅଟେ। ସମସ୍ତ benchmark ସ୍ୱତନ୍ତ୍ର ଭାବରେ ମପାଯାଇଛି।
93.5%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
46.9%
HLE
Humanity's Last Exam
59.5%
SciCode
SciCode - scientific code generation
88.7%
Long Context
Long Context Reasoning - reasoning over long inputs
85.0%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
44.2%
FrontierCode
FrontierCode - long-horizon production coding tasks
60.4%
ARC-AGI-2
ARC-AGI-2 - abstract reasoning on novel puzzles
1674
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
କିମି K3 ବିଷୟରେ
Moonshot AI's most capable model: an open-weight, natively multimodal agentic model with 2.8T total parameters and roughly 104B activated across 896 experts. It is built on Kimi Delta Attention and Attention Residuals with a Stable LatentMoE framework, giving roughly 2.5 times better scaling efficiency than Kimi K2, and pairs native visual understanding with a million-token context. It is designed for frontier work -- software engineering, knowledge work and deep reasoning -- and always reasons.
Moonshot's own framing is that the parameter count is not the point. The architecture changes are what the lab argues for: Kimi Delta Attention as an efficient base for scaling attention, and Attention Residuals, which retrieve representations selectively across depth instead of accumulating them uniformly. Sparsity was pushed further too, with a Stable LatentMoE framework effectively activating 16 of 896 experts, and the routing work that makes that stable at scale - allocation derived directly from router-score quantiles rather than from a hand-tuned balancing term, and an optimiser that treats each attention head independently. The claim Moonshot puts on all of it together is roughly 2.5 times better scaling efficiency than Kimi K2: compute converted into capability, rather than parameters converted into a headline.
The case for the model is made mostly in case studies. Given 24 hours per task in identical sandboxes, Moonshot reports K3 competitive with Claude Fable 5 and substantially ahead of Claude Opus 4.8, GPT 5.6 Sol and GPT 5.5 at optimising GPU kernels across two hardware families, and notes that late in K3's own development an early version handled most of the team's kernel work. Asked to build a GPU programming system from scratch it produced a compact Triton-like compiler with its own tile-level intermediate representation, optimisation passes and a code-generation path down to PTX, which the lab says matches or beats the established stack on supported benchmarks and trains a small transformer end to end. In a single 48-hour autonomous run it designed, optimised and verified a chip for a nano model built on its own architecture, closing timing at 100 MHz inside four square millimetres.
The published table backs that with 88.3 on Terminal-Bench 2.1, 42.0 on SWE Marathon and 77.8 on Program Bench, the highest in its compared set on the last two, alongside 93.5 on GPQA-Diamond and 91.2 on BrowseComp. Multimodality is native rather than bolted on, and Moonshot leans on it for video in particular: it points to K3 cutting its own launch teaser from 56 source clips with beat-synchronised edits, and producing an animated explainer of its own architecture.
The limitations section is the part to read before wiring it into anything. K3 was trained in preserved-thinking mode, so a harness that fails to pass the full thinking history back, or a session handed over to K3 from another model mid-way, can make generation highly unstable. It is also, in the lab's own words, prone to excessive proactiveness: trained hard on long-horizon work, it tends to decide on your behalf when the intent is ambiguous, and needs explicit behavioural constraints in the system prompt if it must stay inside defined boundaries. Moonshot states plainly that overall performance still trails Claude Fable 5 and GPT 5.6 Sol, with a noticeable gap in user experience.
The weights are open. Moonshot committed to publishing the full set by 27 July 2026 and contributed a prefix-caching implementation for its new attention design upstream so the model can be served efficiently outside the lab's own stack. Self-hosting it is a serious undertaking rather than a download, though: the recommendation is deployment on supernode configurations of 64 or more accelerators. The copy served here is Moonshot's.
ଲଞ୍ଚ ସମୟରେ Moonshot AI ଯାହା କହିଥିଲା
- An open 3T-class model
- Moonshot presents K3 as the first open model at 2.8 trillion parameters, and claims roughly 2.5 times better scaling efficiency than Kimi K2 from its new attention and mixture-of-experts design.
- Kernels, compilers, silicon
- The lab reports K3 competitive with Claude Fable 5 at GPU kernel optimisation, building a working Triton-like compiler from scratch, and designing and verifying a small chip in one 48-hour run.
- Long-horizon coding
- On its published table Moonshot reports 88.3 on Terminal-Bench 2.1, plus the top scores in its compared set on SWE Marathon at 42.0 and Program Bench at 77.8.
- Knowledge work end to end
- Native vision is used for research reports with generated charts and interactive narratives, for editing video, and for producing infographic-style presentations rather than only reading images.
- Preserved thinking is required
- K3 was trained with its full thinking history carried forward. A harness that drops it, or a session switched to K3 mid-way from another model, can make output highly unstable.
- What it is not for
- Moonshot warns the model improvises on your behalf when intent is ambiguous and needs explicit constraints, and says overall performance and user experience still trail Claude Fable 5 and GPT 5.6 Sol.
ଭାରତୀୟ ଭାଷା
କିମି K3 7 ଟି ଭାରତୀୟ ଭାଷାରେ ଉତ୍ତର ଦିଏ। ସେହି ଭାଷାରେ ଉତ୍ତର ପାଇବା ପାଇଁ ମେସେଜ୍ ବକ୍ସ ପାଖରେ ଥିବା ମେନୁରୁ ଭାଷା ବାଛନ୍ତୁ।
Frequently Asked Questions
କିମି K3 ବିଷୟରେ ବାରମ୍ବାର ପଚରାଯାଉଥିବା ପ୍ରଶ୍ନ।
କିମି K3 କେବେ ଲଞ୍ଚ ହୋଇଥିଲା?
Moonshot AI ଜୁଲାଇ 16, 2026 ରେ କିମି K3 ଲଞ୍ଚ କରିଥିଲା।
କିମି K3 କିଏ ତିଆରି କରିଛି?
କିମି K3 କୁ Moonshot AI ତିଆରି କରିଛି। 99Models ଏହା ସହ ସିଧାସଳଖ ପ୍ରୋଭାଇଡରଙ୍କ ନିର୍ଦ୍ଧାରିତ ରେଟ୍ ରେ ସଂଯୋଗ କରେ।
କିମି K3 କେତେ ଶକ୍ତିଶାଳୀ?
ଇଣ୍ଟେଲିଜେନ୍ସ ତାଲିକାରେ 54 ଟି chat Model ମଧ୍ୟରୁ ଏହାର ରାଙ୍କ୍ 9। ସ୍ୱାଧୀନ ବେଞ୍ଚମାର୍କ ସ୍କୋର ଆଧାରରେ ଏହା ସ୍ଥିର କରାଯାଇଛି, ଯାହା ଉପରେ ଥିବା ଟେବୁଲରେ ଉପଲବ୍ଧ।
କିମି K3 ର ମୂଲ୍ୟ କେତେ?
ବ୍ୟବହାର ଖର୍ଚ୍ଚ ପ୍ରତି 10 ଲକ୍ଷ ଇନପୁଟ୍ Tokens ପାଇଁ ₹316.80 ଏବଂ ଆଉଟପୁଟ୍ Tokens ପାଇଁ ₹1,584.00, 0% ମାର୍କଅପ୍ ସହିତ। କୌଣସି ସବସ୍କ୍ରିପସନ୍ ନାହିଁ; ଆପଣ ଯେତିକି ବ୍ୟବହାର କରିବେ ସେତିକି ପେମେଣ୍ଟ କରିବେ।
US ଡଲାରରେ କିମି K3 ର ମୂଲ୍ୟ କେତେ?
ପ୍ରୋଭାଇଡର୍ ପ୍ରତି 10 ଲକ୍ଷ ଇନପୁଟ୍ Tokens ପାଇଁ $3.30 ଏବଂ ଆଉଟପୁଟ୍ Tokens ପାଇଁ $16.50 ଚାର୍ଜ କରେ। ଏହି ପୃଷ୍ଠାର ଟଙ୍କା ମୂଲ୍ୟ ₹96 ପ୍ରତି ଡଲାର ହିସାବରେ ରୂପାନ୍ତରିତ।
କିମି K3 କେତେ ଲମ୍ବା କଥାବାର୍ତ୍ତା ମନେ ରଖିପାରିବ?
ଏହାର Context window ହେଉଛି 10.5 lakh Tokens। ଗୋଟିଏ request ରେ ଏହା ସମୁଦାୟ କଥାବାର୍ତ୍ତା ଏବଂ ସଂଲଗ୍ନ ଫାଇଲ୍ ପଢ଼ିପାରିବ।
ମୂଲ୍ୟ ହିସାବରେ କିମି K3 କେତେ ଭଲ?
ଭ୍ୟାଲୁ ରାଙ୍କିଙ୍ଗରେ 54 ଟି Model ମଧ୍ୟରୁ ଏହାର ସ୍ଥାନ 21। ଏହି ରାଙ୍କିଙ୍ଗ ବେଞ୍ଚମାର୍କ କ୍ଷମତା ଏବଂ Token ଖର୍ଚ୍ଚକୁ ତୁଳନା କରି ସ୍ଥିର କରାଯାଏ।
କିମି K3 କ’ଣ Reasoning ସପୋର୍ଟ କରେ?
ଡିଫଲ୍ଟ ଭାବରେ On। ଯେଉଁଠାରେ Reasoning ଉପଲବ୍ଧ, ଆପଣ ମେସେଜ୍ ବକ୍ସରେ thinking effort ସ୍ତର ସେଟ୍ କରିପାରିବେ।
କିମି K3 କେଉଁ ଭାରତୀୟ ଭାଷାରେ ଉତ୍ତର ଦେଇପାରେ?
ଏହା 7 ଟି ଭାରତୀୟ ଭାଷାରେ ଉତ୍ତର ଦିଏ। ମେସେଜ୍ ବକ୍ସ ପାଖରେ ଥିବା ମେନୁରୁ ଆପଣଙ୍କ ପସନ୍ଦର ଭାଷା ବାଛନ୍ତୁ।
Moonshot AI ରୁ ଅନ୍ୟାନ୍ୟ Models
କାଟାଲଗ୍ ଅପଡେଟ୍ ହୋଇଛି: ସେପ୍ଟେମ୍ବର 9, 2026