99MODELS

কিউৱেন 3 ক'ডাৰ নেক্সট

Cheap coding-specialised Qwen for repetitive code edits.

মুকলিৰ তাৰিখ: 4 ফেব্ৰু 2026

76.80

প্ৰতি 10 লাখ আউটপুট Tokens

ইনপুট: প্ৰতি 10 লাখ Tokens-ত ₹11.52

প্ৰতি মাৰ্কিন ডলাৰত ₹96 হাৰত ৰূপান্তৰিত, 0% মাৰ্কআপৰ সৈতে প্ৰভাইডাৰৰ নিৰ্ধাৰিত দৰত বিল কৰা হয়।

Specifications

Context window
2,62,144 tokens
সৰ্বাধিক আউটপুট
2,35,929 tokens
গ্ৰহণ কৰে
টেক্সট
Reasoning
নহয়
টুলৰ ব্যৱহাৰ
হয়
গঠনবদ্ধ আউটপুট
হয়
ক'ড কাৰ্যকৰীকৰণ
নহয়
বুদ্ধিমত্তাৰ ৰেংক
54 ৰ ভিতৰত #49
ভ্যালু ৰেংক
54 ৰ ভিতৰত #45

মূল্য

মূল্য
প্ৰতি 10 lakh TokensINRUSD
ইনপুট11.52$0.12
আউটপুট76.80$0.80
কেশ্বড ইনপুট6.72$0.07

Benchmarks

ৰেটিং হিচাপে উল্লেখ নথকালৈকে স্ক’ৰসমূহ শতাংশত দিয়া হৈছে। সকলো Benchmark স্বতন্ত্ৰভাৱে পৰীক্ষা কৰা হৈছে।

  • 73.7%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 10.1%

    HLE

    Humanity's Last Exam

  • 36.2%

    SciCode

    SciCode - scientific code generation

  • 35.2%

    IFBench

    IFBench - precise instruction following

  • 47.0%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 18.2%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 38.2%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

কিউৱেন 3 ক'ডাৰ নেক্সট-ৰ বিষয়ে

Qwen's coding model for agents and local development, optimised for repository-level understanding and multi-turn tool interaction. With only 3B parameters activated out of 80B total it reaches performance comparable to models with ten to twenty times more active parameters, which is what makes it cheap enough to run in a loop. It excels at long-horizon reasoning, complex tool use and recovery from execution failures, and it does not reason.

Alibaba's stated approach here was to scale the training signal rather than the parameter count. The model is built on an existing hybrid-attention sparse base and then trained agentically at scale: continued pretraining on code- and agent-centric data, supervised fine-tuning on high-quality agent trajectories, domain-specialised expert training in software engineering, QA and web or UX work, and finally distillation of those experts back into one deployment-ready model. The training tasks were paired with executable environments throughout, so the model learned from what actually happened when it ran something rather than from descriptions of correct code.

That recipe is aimed at three specific behaviours, and Alibaba names them: long-horizon reasoning, tool use, and recovery from execution failures. The third is the one that decides whether a coding agent is usable. A model that writes a good patch but cannot read a stack trace and try again needs a human at every step; one that can recover turns a failed run into the next attempt on its own.

The lab reports over 70% on SWE-Bench Verified with an agent scaffold, competitive results on the multilingual variant and on the harder SWE-Bench Pro, and - the more interesting claim - that its SWE-Bench Pro score keeps climbing as the agent is allowed more turns. Alibaba offers that as evidence the model holds up over long multi-turn runs rather than peaking on a single shot. With only 3 billion parameters active out of 80 billion, it says the model matches or exceeds several much larger open models on agent-centric evaluations, which is what puts it on a useful efficiency frontier rather than merely at a low price.

The economics are the practical reason to reach for it: at this rate, letting the agent retry is not a budget decision. Repository-level edits, running tests, reading the failure and going round again is exactly the loop it was trained in, and the lab demonstrates it inside terminal agents, editor extensions and browser agents rather than in chat.

Be clear about what it is not. It does not reason - there is no thinking pass and no effort dial - and it is a coding specialist, not a general assistant: on general knowledge and reasoning evaluations it lands far below any of the lab's flagship models. Alibaba's own summary says there is still much room for improvement and names tool use, hard problems and complex task management as the things it wants to fix next. Use it for repetitive, verifiable code work where a test can say whether it succeeded; take a reasoning model for architecture decisions, novel algorithms, or anything that is not code.

মুকলিৰ সময়ত Alibaba-এ যি কৈছিল

Agentic training, not scale
Alibaba scaled verifiable coding tasks paired with executable environments instead of scaling parameters, so the model learned from what happened when its code ran.
Three billion active parameters
Only 3 billion of its 80 billion parameters activate per token, which the lab says matches or exceeds open models with ten to twenty times more active parameters on agent evaluations.
Holds up over many turns
Alibaba reports its SWE-Bench Pro score rising as the agent is given more turns, offered as evidence that it sustains long multi-turn work rather than peaking on one attempt.
Recovery from failures
The training recipe explicitly targets recovery from execution failures, which is what lets a coding agent read a failing test and try again without a human in the loop.
No reasoning pass
There is no thinking step and no effort dial. Responses come straight out, which is part of why it is cheap enough to run in a loop.
What it is not for
It is a coding specialist and sits far below the flagship models on general reasoning. The lab itself names tool use, hard problems and complex task management as work still to do.

Frequently Asked Questions

কিউৱেন 3 ক'ডাৰ নেক্সট সম্পৰ্কে সঘনাই সোধা প্ৰশ্নসমূহ।

কিউৱেন 3 ক'ডাৰ নেক্সট কেতিয়া মুকলি কৰা হৈছিল?

Alibaba-এ কিউৱেন 3 ক'ডাৰ নেক্সট মডেলটো 4 ফেব্ৰু 2026 তাৰিখে মুকলি কৰিছিল।

কিউৱেন 3 ক'ডাৰ নেক্সট কোনে তৈয়াৰ কৰিছে?

কিউৱেন 3 ক'ডাৰ নেক্সট-ক AI লেব Alibaba-এ নিৰ্মাণ কৰিছে। 99Models AI-য়ে প্ৰভাইডাৰৰ নিৰ্ধাৰিত দৰতে ইয়াৰ পোনপটীয়া সংযোগ প্ৰদান কৰে।

কিউৱেন 3 ক'ডাৰ নেক্সট কিমান বুদ্ধিমান?

Intelligence তালিকাত 54 টা Chat Model-ৰ ভিতৰত ইয়াৰ স্থান 49, যিটো স্বতন্ত্ৰ Benchmark স্কোৰৰ দ্বাৰা নিৰ্ধাৰিত। ওপৰৰ Benchmarks তালিকাত ইয়াৰ সম্পূৰ্ণ স্কোৰ উপলব্ধ।

কিউৱেন 3 ক'ডাৰ নেক্সট-ৰ খৰচ কিমান?

প্ৰতি 10 লাখ Input Tokens-ত ₹11.52 আৰু Output Tokens-ত ₹76.80, 0% মাৰ্কআপসহ প্ৰভাইডাৰৰ দৰত চাৰ্জ কৰা হয়। কোনো চাবস্ক্ৰিপশ্বন নাই; ব্যৱহাৰ অনুসৰি পেমেন্ট কৰক।

মাৰ্কিন ডলাৰত কিউৱেন 3 ক'ডাৰ নেক্সট-ৰ API মূল্য কিমান?

প্ৰভাইডাৰে প্ৰতি 10 লাখ Input Tokens-ত $0.12 আৰু Output Tokens-ত $0.80 চাৰ্জ কৰে। ₹96 ডলাৰ বিনিময় হাৰত টকালৈ ৰূপান্তৰ কৰা হৈছে।

কিউৱেন 3 ক'ডাৰ নেক্সট-এ কিমান দীঘলীয়া কথা-বতৰা মনত ৰাখিব পাৰে?

ইয়াৰ Context Window হ'ল 2.6 lakh Tokens। এটা ৰিকুৱেষ্টত এতিয়ালৈকে হোৱা কথোপকথন আৰু সংলগ্ন ফাইলসমূহ ই একেলগে প্ৰক্ৰিয়াকৰণ কৰিব পাৰে।

মূল্য আৰু পাৰদৰ্শিতাৰ ফালৰ পৰা কিউৱেন 3 ক'ডাৰ নেক্সট কিমান লাভজনক?

Value তালিকাত 54 টা Model-ৰ ভিতৰত ইয়াৰ স্থান 45। এই ৰেংকিং Benchmark বুদ্ধিমত্তা আৰু প্ৰকৃত Tokens খৰচৰ তুলনা কৰি নিৰ্ধাৰণ কৰা হয়।

Alibaba-ৰ অন্যান্য Models

কেটেলগ আপডেট কৰা হৈছে 9 ছেপ্তে 2026