99MODELS

क्वेन 3 कोडर नेक्स्ट

Cheap coding-specialised Qwen for repetitive code edits.

रिलिज मिति: 2026 फेब्रुअरी 4

76.80

प्रति 10 लाख आउटपुट Tokens

इनपुट: ₹11.52 प्रति 10 लाख Tokens

प्रति अमेरिकी डलर ₹96 मा 0% मार्कअपका साथ प्रदायककै दरमा गणना गरिन्छ।

विवरण

Context विन्डो
2,62,144 Tokens
अधिकतम आउटपुट
2,35,929 Tokens
स्वीकार गर्छ
टेक्स्ट
Reasoning
छैन
टुल प्रयोग
संरचित आउटपुट
कोड कार्यान्वयन
छैन
इन्टेलिजेन्स र्‍याङ्क
54 मध्ये #49
भ्याल्यू र्‍याङ्क
54 मध्ये #45

मूल्य

मूल्य
प्रति 10 lakh TokensINRUSD
इनपुट11.52$0.12
आउटपुट76.80$0.80
क्यास गरिएको इनपुट6.72$0.07

बेन्चमार्क

रेटिङ बाहेकका सबै स्कोरहरू प्रतिशतमा छन्। सबै बेन्चमार्कहरू स्वतन्त्र रूपमा मापन गरिएका हुन्।

  • 73.7%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 10.1%

    HLE

    Humanity's Last Exam

  • 36.2%

    SciCode

    SciCode - scientific code generation

  • 35.2%

    IFBench

    IFBench - precise instruction following

  • 47.0%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 18.2%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 38.2%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

क्वेन 3 कोडर नेक्स्ट को बारेमा

Qwen's coding model for agents and local development, optimised for repository-level understanding and multi-turn tool interaction. With only 3B parameters activated out of 80B total it reaches performance comparable to models with ten to twenty times more active parameters, which is what makes it cheap enough to run in a loop. It excels at long-horizon reasoning, complex tool use and recovery from execution failures, and it does not reason.

Alibaba's stated approach here was to scale the training signal rather than the parameter count. The model is built on an existing hybrid-attention sparse base and then trained agentically at scale: continued pretraining on code- and agent-centric data, supervised fine-tuning on high-quality agent trajectories, domain-specialised expert training in software engineering, QA and web or UX work, and finally distillation of those experts back into one deployment-ready model. The training tasks were paired with executable environments throughout, so the model learned from what actually happened when it ran something rather than from descriptions of correct code.

That recipe is aimed at three specific behaviours, and Alibaba names them: long-horizon reasoning, tool use, and recovery from execution failures. The third is the one that decides whether a coding agent is usable. A model that writes a good patch but cannot read a stack trace and try again needs a human at every step; one that can recover turns a failed run into the next attempt on its own.

The lab reports over 70% on SWE-Bench Verified with an agent scaffold, competitive results on the multilingual variant and on the harder SWE-Bench Pro, and - the more interesting claim - that its SWE-Bench Pro score keeps climbing as the agent is allowed more turns. Alibaba offers that as evidence the model holds up over long multi-turn runs rather than peaking on a single shot. With only 3 billion parameters active out of 80 billion, it says the model matches or exceeds several much larger open models on agent-centric evaluations, which is what puts it on a useful efficiency frontier rather than merely at a low price.

The economics are the practical reason to reach for it: at this rate, letting the agent retry is not a budget decision. Repository-level edits, running tests, reading the failure and going round again is exactly the loop it was trained in, and the lab demonstrates it inside terminal agents, editor extensions and browser agents rather than in chat.

Be clear about what it is not. It does not reason - there is no thinking pass and no effort dial - and it is a coding specialist, not a general assistant: on general knowledge and reasoning evaluations it lands far below any of the lab's flagship models. Alibaba's own summary says there is still much room for improvement and names tool use, hard problems and complex task management as the things it wants to fix next. Use it for repetitive, verifiable code work where a test can say whether it succeeded; take a reasoning model for architecture decisions, novel algorithms, or anything that is not code.

लन्चको समयमा Alibaba ले के भन्यो

Agentic training, not scale
Alibaba scaled verifiable coding tasks paired with executable environments instead of scaling parameters, so the model learned from what happened when its code ran.
Three billion active parameters
Only 3 billion of its 80 billion parameters activate per token, which the lab says matches or exceeds open models with ten to twenty times more active parameters on agent evaluations.
Holds up over many turns
Alibaba reports its SWE-Bench Pro score rising as the agent is given more turns, offered as evidence that it sustains long multi-turn work rather than peaking on one attempt.
Recovery from failures
The training recipe explicitly targets recovery from execution failures, which is what lets a coding agent read a failing test and try again without a human in the loop.
No reasoning pass
There is no thinking step and no effort dial. Responses come straight out, which is part of why it is cheap enough to run in a loop.
What it is not for
It is a coding specialist and sits far below the flagship models on general reasoning. The lab itself names tool use, hard problems and complex task management as work still to do.

Frequently Asked Questions

क्वेन 3 कोडर नेक्स्ट सम्बन्धी प्रायः सोधिने प्रश्नहरू।

क्वेन 3 कोडर नेक्स्ट कहिले रिलिज भएको हो?

Alibaba ले क्वेन 3 कोडर नेक्स्ट लाई 2026 फेब्रुअरी 4 मा सार्वजनिक गरेको हो।

क्वेन 3 कोडर नेक्स्ट कसले बनाएको हो?

क्वेन 3 कोडर नेक्स्ट लाई Alibaba ले बनाएको हो। 99Models AI ले प्रदायककै दरमा सिधै जोड्दछ।

क्वेन 3 कोडर नेक्स्ट कत्तिको सक्षम र बुद्धिमानी छ?

यो हाम्रो बौद्धिकता श्रेणीकरणमा 54 च्याट Models मध्ये 49 स्थानमा छ। यसको पूर्ण अङ्क माथिको Benchmarks तालिकामा हेर्न सकिन्छ।

क्वेन 3 कोडर नेक्स्ट को लागत कति पर्छ?

यसमा 0% मार्कअपका साथ प्रति 10 लाख इनपुट Tokens को ₹11.52 र आउटपुटको ₹76.80 लाग्छ। कुनै सदस्यता छैन; तपाईंले प्रयोग गरेअनुसार मात्र भुक्तानी गर्नुहुन्छ।

डलरमा क्वेन 3 कोडर नेक्स्ट को API मूल्य कति हो?

प्रदायकले प्रति 10 लाख इनपुट Tokens को $0.12 र आउटपुट Tokens को $0.80 शुल्क लिन्छ। यस पृष्ठका दरहरू प्रति अमेरिकी डलर ₹96 मा रूपान्तरण गरिएका हुन्।

क्वेन 3 कोडर नेक्स्ट ले कति लामो कुराकानी सम्झन सक्छ?

यसको Context विन्डो 2.6 lakh Tokens हो। यसले एकल अनुरोधमा प्रक्रिया गर्न सक्ने कुराकानी र संलग्न फाइलहरूको कुल क्षमता यही हो।

के क्वेन 3 कोडर नेक्स्ट लागत अनुसार उत्कृष्ट छ?

मूल्य र गुणस्तरको आधारमा यो 54 Models मध्ये 45 स्थानमा छ। यसले बौद्धिकता र Token लागतको तुलना गर्दछ।

Alibaba का अन्य Models

क्याटलग अद्यावधिक गरिएको मिति: 2026 सेप्टेम्बर 9