کیوین 3.8 میکس
نیاAlibaba-hosted 3.8 flagship; 256K window on our one route.
تاریخ اجراء: 3 ستمبر، 2026
₹576.00
فی 10 لاکھ آؤٹ پٹ Tokens
ان پٹ: ₹192.00 فی 10 لاکھ Tokens
فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔
تکنیکی تفصیلات
- Context ونڈو
- 2,56,000 Tokens
- زیادہ سے زیادہ آؤٹ پٹ
- 1,31,072 Tokens
- سپورٹ
- ٹیکسٹ
- Reasoning
- پہلے سے آن
- ایفرٹ لیولز
- minimal, low, medium, high, xhigh
- ٹول کا استعمال
- ہاں
- سٹرکچرڈ آؤٹ پٹ
- ہاں
- کوڈ ایگزیکیوشن
- نہیں
- پیرامیٹرز
- 2.4T MoE
- ذہانت کا درجہ
- 54 میں سے #13
- ویلیو کا درجہ
- 54 میں سے #14
قیمتیں
Benchmarks
تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔
92.7%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
43.0%
HLE
Humanity's Last Exam
53.2%
SciCode
SciCode - scientific code generation
78.3%
Long Context
Long Context Reasoning - reasoning over long inputs
81.3%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
کیوین 3.8 میکس کے بارے میں
Alibaba's proprietary 3.8 flagship: a 2.4-trillion-parameter sparse Mixture-of-Experts model positioned as a comprehensive step up in coding and professional knowledge work over the previous Max generation. It is the hosted, closed counterpart to the openly released 3.8 2.4T, and it always reasons. Note that the window here is 256K rather than the million tokens the model is capable of, because the one host we route to serves it at that length.
Alibaba argued this model with long autonomous runs rather than with single scores, and the three coding cases in its launch post are all days long. Asked to create a command-line project from an empty folder, the model ran for more than ten days building a harness that folds community feedback, its own test results and normalised issues into one loop; after roughly sixteen days of unattended operation the lab counted 265 commits, 127 pull requests and 151 issues in the repository. Given a research paper, a set of GPUs and no starter code, it worked about 125 hours, wrote roughly 7,600 lines and ran 33 rounds of GPU training, reproducing the paper's six findings in the first 37 hours and then inventing and testing 18 ideas of its own across four rounds until it beat the paper's own method by 2.7 points on AIME24. Entered into a live online contest against 526 human teams under a 24-hour limit, it climbed from 0.60 to 0.853 accuracy over 45 submissions and finished ahead of 458 of them.
The thread Alibaba draws through all three is that the model does not just execute a plan - it revises the plan from feedback. That is also how the lab describes its training for ordinary work: reinforcement-learning environments and compute scaled together along task, workspace and harness axes, with one reward system spanning execution checks, rubric judging of both text and rendered visual output, and agentic inspection, plus a batch balancer to keep training stable as it scales. The stated aim is competence that lifts across harnesses rather than in one, and Alibaba reports comparable results across several popular agent harnesses.
It then stress-tested breadth across several hundred high-value professions. The examples the lab publishes are concrete: a compliance review that surfaced 1,284 relevant clauses across hundreds of documents in under an hour, an eight-screen banking prototype with a consistent design system delivered in one shot with no revision rounds, a 26-dish menu costed to a 33.8% food-cost ratio from a hundred-odd supply briefs, and a seismic model of a 30-storey tower reconstructed in the browser from one set of drawings.
For long-horizon work Alibaba reports two results worth quoting. On a 365-day e-commerce operations simulation built from real transaction data, with nearly 600 suppliers of which 152 were deliberately fraudulent, the model ended the year with the highest balance of the field, 416,252 yuan from 100,000 in starting capital - 38% ahead of the next model and 152% above the previous Max generation. On an autonomous chip-design task it took a cryptographic accelerator from 8,298 gates to 678 across roughly 500 turns, with the physical die shrinking 81% after place and route, and Alibaba notes the largest single structural improvement came hundreds of turns into the run rather than early.
On its own benchmark table Alibaba reports the best score in the comparison on PaperBench at 93.0, WideSearch at 81.9, IFBench at 82.8 and HealthBench at 60.2, alongside 86.6 on Terminal Bench 2.1, 92.6 on GPQA Diamond and 43.6 on HLE. It also publishes where it trails: 56.6 on DeepSWE v1.1 against 73.0 and 70.0 for the two models it compares with, and 67.7 on SWE-bench Pro against 80.0. Reasoning is always on and the effort dial defaults to its highest setting, with reasoning from earlier turns preserved by default; the lab's own caution is that on multi-turn agentic work a lower effort setting does not reliably finish sooner, because thinner analysis produces more failures and retries.
لانچ کے وقت Alibaba کا بیان
- Ten days unattended
- Alibaba reports a single autonomous run of more than ten days that produced 265 commits, 127 pull requests and 151 issues while the model maintained and extended its own harness.
- Reproducing and improving a paper
- From nothing but a paper and GPUs, the lab says it rebuilt the full pipeline in about 37 hours, then tested 18 ideas of its own and beat the paper's method by 2.7 points on AIME24.
- Work across professions
- Alibaba publishes examples from several hundred high-value professions, including 1,284 contract clauses surfaced in under an hour and an eight-screen prototype delivered with no revision rounds.
- Long-horizon operations
- On a 365-day business simulation the lab reports the highest final balance in the field, 38% ahead of the next model, and on a chip-design task a reduction from 8,298 gates to 678.
- Benchmarks the lab leads
- Alibaba reports the best score in its comparison on PaperBench at 93.0, WideSearch at 81.9, IFBench at 82.8 and HealthBench at 60.2, with 92.6 on GPQA Diamond.
- What it is not for
- The same table has it behind on DeepSWE v1.1 at 56.6 against 73.0, and on SWE-bench Pro at 67.7 against 80.0. Reasoning cannot be disabled, so there is no fast non-thinking mode.
Frequently Asked Questions
کیوین 3.8 میکس کے بارے میں اکثر پوچھے جانے والے سوالات۔
کیوین 3.8 میکس کب جاری ہوا تھا؟
Alibaba نے کیوین 3.8 میکس کو 3 ستمبر، 2026 کو جاری کیا۔
کیوین 3.8 میکس کس نے بنایا ہے؟
کیوین 3.8 میکس کو Alibaba نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔
کیوین 3.8 میکس کتنا ذہین ہے؟
یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 13 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔
کیوین 3.8 میکس کا کتنا خرچ آتا ہے؟
استعمال کی لاگت ₹192.00 فی 10 لاکھ ان پٹ Tokens اور ₹576.00 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔
امریکی ڈالر میں کیوین 3.8 میکس کی قیمت کیا ہے؟
فراہم کنندہ $2.00 فی 10 لاکھ ان پٹ Tokens اور $6.00 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔
کیوین 3.8 میکس کتنی لمبی گفتگو یاد رکھ سکتا ہے؟
اس کی Context حد 2.6 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔
کیا کیوین 3.8 میکس مناسب قیمت میں بہترین کارکردگی دیتا ہے؟
یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 14 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔
کیا کیوین 3.8 میکس جواب دینے سے پہلے سوچتا ہے؟
پہلے سے آن۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔
Alibaba کے مزید ماڈلز
کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026