99MODELS

کیوین 3 کوڈر نیکسٹ

Cheap coding-specialised Qwen for repetitive code edits.

تاریخ اجراء: 4 فروری، 2026

76.80

فی 10 لاکھ آؤٹ پٹ Tokens

ان پٹ: ₹11.52 فی 10 لاکھ Tokens

فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔

تکنیکی تفصیلات

Context ونڈو
2,62,144 Tokens
زیادہ سے زیادہ آؤٹ پٹ
2,35,929 Tokens
سپورٹ
ٹیکسٹ
Reasoning
نہیں
ٹول کا استعمال
ہاں
سٹرکچرڈ آؤٹ پٹ
ہاں
کوڈ ایگزیکیوشن
نہیں
ذہانت کا درجہ
54 میں سے #49
ویلیو کا درجہ
54 میں سے #45

قیمتیں

قیمتیں
فی 10 lakh TokensINRUSD
ان پٹ11.52$0.12
آؤٹ پٹ76.80$0.80
کیشڈ ان پٹ6.72$0.07

Benchmarks

تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔

  • 73.7%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 10.1%

    HLE

    Humanity's Last Exam

  • 36.2%

    SciCode

    SciCode - scientific code generation

  • 35.2%

    IFBench

    IFBench - precise instruction following

  • 47.0%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 18.2%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 38.2%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

کیوین 3 کوڈر نیکسٹ کے بارے میں

Qwen's coding model for agents and local development, optimised for repository-level understanding and multi-turn tool interaction. With only 3B parameters activated out of 80B total it reaches performance comparable to models with ten to twenty times more active parameters, which is what makes it cheap enough to run in a loop. It excels at long-horizon reasoning, complex tool use and recovery from execution failures, and it does not reason.

Alibaba's stated approach here was to scale the training signal rather than the parameter count. The model is built on an existing hybrid-attention sparse base and then trained agentically at scale: continued pretraining on code- and agent-centric data, supervised fine-tuning on high-quality agent trajectories, domain-specialised expert training in software engineering, QA and web or UX work, and finally distillation of those experts back into one deployment-ready model. The training tasks were paired with executable environments throughout, so the model learned from what actually happened when it ran something rather than from descriptions of correct code.

That recipe is aimed at three specific behaviours, and Alibaba names them: long-horizon reasoning, tool use, and recovery from execution failures. The third is the one that decides whether a coding agent is usable. A model that writes a good patch but cannot read a stack trace and try again needs a human at every step; one that can recover turns a failed run into the next attempt on its own.

The lab reports over 70% on SWE-Bench Verified with an agent scaffold, competitive results on the multilingual variant and on the harder SWE-Bench Pro, and - the more interesting claim - that its SWE-Bench Pro score keeps climbing as the agent is allowed more turns. Alibaba offers that as evidence the model holds up over long multi-turn runs rather than peaking on a single shot. With only 3 billion parameters active out of 80 billion, it says the model matches or exceeds several much larger open models on agent-centric evaluations, which is what puts it on a useful efficiency frontier rather than merely at a low price.

The economics are the practical reason to reach for it: at this rate, letting the agent retry is not a budget decision. Repository-level edits, running tests, reading the failure and going round again is exactly the loop it was trained in, and the lab demonstrates it inside terminal agents, editor extensions and browser agents rather than in chat.

Be clear about what it is not. It does not reason - there is no thinking pass and no effort dial - and it is a coding specialist, not a general assistant: on general knowledge and reasoning evaluations it lands far below any of the lab's flagship models. Alibaba's own summary says there is still much room for improvement and names tool use, hard problems and complex task management as the things it wants to fix next. Use it for repetitive, verifiable code work where a test can say whether it succeeded; take a reasoning model for architecture decisions, novel algorithms, or anything that is not code.

لانچ کے وقت Alibaba کا بیان

Agentic training, not scale
Alibaba scaled verifiable coding tasks paired with executable environments instead of scaling parameters, so the model learned from what happened when its code ran.
Three billion active parameters
Only 3 billion of its 80 billion parameters activate per token, which the lab says matches or exceeds open models with ten to twenty times more active parameters on agent evaluations.
Holds up over many turns
Alibaba reports its SWE-Bench Pro score rising as the agent is given more turns, offered as evidence that it sustains long multi-turn work rather than peaking on one attempt.
Recovery from failures
The training recipe explicitly targets recovery from execution failures, which is what lets a coding agent read a failing test and try again without a human in the loop.
No reasoning pass
There is no thinking step and no effort dial. Responses come straight out, which is part of why it is cheap enough to run in a loop.
What it is not for
It is a coding specialist and sits far below the flagship models on general reasoning. The lab itself names tool use, hard problems and complex task management as work still to do.

Frequently Asked Questions

کیوین 3 کوڈر نیکسٹ کے بارے میں اکثر پوچھے جانے والے سوالات۔

کیوین 3 کوڈر نیکسٹ کب جاری ہوا تھا؟

Alibaba نے کیوین 3 کوڈر نیکسٹ کو 4 فروری، 2026 کو جاری کیا۔

کیوین 3 کوڈر نیکسٹ کس نے بنایا ہے؟

کیوین 3 کوڈر نیکسٹ کو Alibaba نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔

کیوین 3 کوڈر نیکسٹ کتنا ذہین ہے؟

یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 49 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔

کیوین 3 کوڈر نیکسٹ کا کتنا خرچ آتا ہے؟

استعمال کی لاگت ₹11.52 فی 10 لاکھ ان پٹ Tokens اور ₹76.80 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔

امریکی ڈالر میں کیوین 3 کوڈر نیکسٹ کی قیمت کیا ہے؟

فراہم کنندہ $0.12 فی 10 لاکھ ان پٹ Tokens اور $0.80 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔

کیوین 3 کوڈر نیکسٹ کتنی لمبی گفتگو یاد رکھ سکتا ہے؟

اس کی Context حد 2.6 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔

کیا کیوین 3 کوڈر نیکسٹ مناسب قیمت میں بہترین کارکردگی دیتا ہے؟

یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 45 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔

Alibaba کے مزید ماڈلز

کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026