99MODELS

क्वेन 3.8 म्याक्स

नयाँ

Alibaba-hosted 3.8 flagship; 256K window on our one route.

रिलिज मिति: 2026 सेप्टेम्बर 3

576.00

प्रति 10 लाख आउटपुट Tokens

इनपुट: ₹192.00 प्रति 10 लाख Tokens

प्रति अमेरिकी डलर ₹96 मा 0% मार्कअपका साथ प्रदायककै दरमा गणना गरिन्छ।

विवरण

Context विन्डो
2,56,000 Tokens
अधिकतम आउटपुट
1,31,072 Tokens
स्वीकार गर्छ
टेक्स्ट
Reasoning
सुरुमै चालु
प्रयास स्तर
minimal, low, medium, high, xhigh
टुल प्रयोग
संरचित आउटपुट
कोड कार्यान्वयन
छैन
प्यारामिटरहरू
2.4T MoE
इन्टेलिजेन्स र्‍याङ्क
54 मध्ये #13
भ्याल्यू र्‍याङ्क
54 मध्ये #14

मूल्य

मूल्य
प्रति 10 lakh TokensINRUSD
इनपुट192.00$2.00
आउटपुट576.00$6.00
क्यास गरिएको इनपुट24.00$0.25

बेन्चमार्क

रेटिङ बाहेकका सबै स्कोरहरू प्रतिशतमा छन्। सबै बेन्चमार्कहरू स्वतन्त्र रूपमा मापन गरिएका हुन्।

  • 92.7%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 43.0%

    HLE

    Humanity's Last Exam

  • 53.2%

    SciCode

    SciCode - scientific code generation

  • 78.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 81.3%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

क्वेन 3.8 म्याक्स को बारेमा

Alibaba's proprietary 3.8 flagship: a 2.4-trillion-parameter sparse Mixture-of-Experts model positioned as a comprehensive step up in coding and professional knowledge work over the previous Max generation. It is the hosted, closed counterpart to the openly released 3.8 2.4T, and it always reasons. Note that the window here is 256K rather than the million tokens the model is capable of, because the one host we route to serves it at that length.

Alibaba argued this model with long autonomous runs rather than with single scores, and the three coding cases in its launch post are all days long. Asked to create a command-line project from an empty folder, the model ran for more than ten days building a harness that folds community feedback, its own test results and normalised issues into one loop; after roughly sixteen days of unattended operation the lab counted 265 commits, 127 pull requests and 151 issues in the repository. Given a research paper, a set of GPUs and no starter code, it worked about 125 hours, wrote roughly 7,600 lines and ran 33 rounds of GPU training, reproducing the paper's six findings in the first 37 hours and then inventing and testing 18 ideas of its own across four rounds until it beat the paper's own method by 2.7 points on AIME24. Entered into a live online contest against 526 human teams under a 24-hour limit, it climbed from 0.60 to 0.853 accuracy over 45 submissions and finished ahead of 458 of them.

The thread Alibaba draws through all three is that the model does not just execute a plan - it revises the plan from feedback. That is also how the lab describes its training for ordinary work: reinforcement-learning environments and compute scaled together along task, workspace and harness axes, with one reward system spanning execution checks, rubric judging of both text and rendered visual output, and agentic inspection, plus a batch balancer to keep training stable as it scales. The stated aim is competence that lifts across harnesses rather than in one, and Alibaba reports comparable results across several popular agent harnesses.

It then stress-tested breadth across several hundred high-value professions. The examples the lab publishes are concrete: a compliance review that surfaced 1,284 relevant clauses across hundreds of documents in under an hour, an eight-screen banking prototype with a consistent design system delivered in one shot with no revision rounds, a 26-dish menu costed to a 33.8% food-cost ratio from a hundred-odd supply briefs, and a seismic model of a 30-storey tower reconstructed in the browser from one set of drawings.

For long-horizon work Alibaba reports two results worth quoting. On a 365-day e-commerce operations simulation built from real transaction data, with nearly 600 suppliers of which 152 were deliberately fraudulent, the model ended the year with the highest balance of the field, 416,252 yuan from 100,000 in starting capital - 38% ahead of the next model and 152% above the previous Max generation. On an autonomous chip-design task it took a cryptographic accelerator from 8,298 gates to 678 across roughly 500 turns, with the physical die shrinking 81% after place and route, and Alibaba notes the largest single structural improvement came hundreds of turns into the run rather than early.

On its own benchmark table Alibaba reports the best score in the comparison on PaperBench at 93.0, WideSearch at 81.9, IFBench at 82.8 and HealthBench at 60.2, alongside 86.6 on Terminal Bench 2.1, 92.6 on GPQA Diamond and 43.6 on HLE. It also publishes where it trails: 56.6 on DeepSWE v1.1 against 73.0 and 70.0 for the two models it compares with, and 67.7 on SWE-bench Pro against 80.0. Reasoning is always on and the effort dial defaults to its highest setting, with reasoning from earlier turns preserved by default; the lab's own caution is that on multi-turn agentic work a lower effort setting does not reliably finish sooner, because thinner analysis produces more failures and retries.

लन्चको समयमा Alibaba ले के भन्यो

Ten days unattended
Alibaba reports a single autonomous run of more than ten days that produced 265 commits, 127 pull requests and 151 issues while the model maintained and extended its own harness.
Reproducing and improving a paper
From nothing but a paper and GPUs, the lab says it rebuilt the full pipeline in about 37 hours, then tested 18 ideas of its own and beat the paper's method by 2.7 points on AIME24.
Work across professions
Alibaba publishes examples from several hundred high-value professions, including 1,284 contract clauses surfaced in under an hour and an eight-screen prototype delivered with no revision rounds.
Long-horizon operations
On a 365-day business simulation the lab reports the highest final balance in the field, 38% ahead of the next model, and on a chip-design task a reduction from 8,298 gates to 678.
Benchmarks the lab leads
Alibaba reports the best score in its comparison on PaperBench at 93.0, WideSearch at 81.9, IFBench at 82.8 and HealthBench at 60.2, with 92.6 on GPQA Diamond.
What it is not for
The same table has it behind on DeepSWE v1.1 at 56.6 against 73.0, and on SWE-bench Pro at 67.7 against 80.0. Reasoning cannot be disabled, so there is no fast non-thinking mode.

Frequently Asked Questions

क्वेन 3.8 म्याक्स सम्बन्धी प्रायः सोधिने प्रश्नहरू।

क्वेन 3.8 म्याक्स कहिले रिलिज भएको हो?

Alibaba ले क्वेन 3.8 म्याक्स लाई 2026 सेप्टेम्बर 3 मा सार्वजनिक गरेको हो।

क्वेन 3.8 म्याक्स कसले बनाएको हो?

क्वेन 3.8 म्याक्स लाई Alibaba ले बनाएको हो। 99Models AI ले प्रदायककै दरमा सिधै जोड्दछ।

क्वेन 3.8 म्याक्स कत्तिको सक्षम र बुद्धिमानी छ?

यो हाम्रो बौद्धिकता श्रेणीकरणमा 54 च्याट Models मध्ये 13 स्थानमा छ। यसको पूर्ण अङ्क माथिको Benchmarks तालिकामा हेर्न सकिन्छ।

क्वेन 3.8 म्याक्स को लागत कति पर्छ?

यसमा 0% मार्कअपका साथ प्रति 10 लाख इनपुट Tokens को ₹192.00 र आउटपुटको ₹576.00 लाग्छ। कुनै सदस्यता छैन; तपाईंले प्रयोग गरेअनुसार मात्र भुक्तानी गर्नुहुन्छ।

डलरमा क्वेन 3.8 म्याक्स को API मूल्य कति हो?

प्रदायकले प्रति 10 लाख इनपुट Tokens को $2.00 र आउटपुट Tokens को $6.00 शुल्क लिन्छ। यस पृष्ठका दरहरू प्रति अमेरिकी डलर ₹96 मा रूपान्तरण गरिएका हुन्।

क्वेन 3.8 म्याक्स ले कति लामो कुराकानी सम्झन सक्छ?

यसको Context विन्डो 2.6 lakh Tokens हो। यसले एकल अनुरोधमा प्रक्रिया गर्न सक्ने कुराकानी र संलग्न फाइलहरूको कुल क्षमता यही हो।

के क्वेन 3.8 म्याक्स लागत अनुसार उत्कृष्ट छ?

मूल्य र गुणस्तरको आधारमा यो 54 Models मध्ये 14 स्थानमा छ। यसले बौद्धिकता र Token लागतको तुलना गर्दछ।

के क्वेन 3.8 म्याक्स ले जवाफ दिनुअघि विचार गर्छ?

सुरुमै चालु। Reasoning उपलब्ध भएको ठाउँमा तपाईंले सिधै सन्देश बक्समा सोच्ने क्षमता समायोजन गर्न सक्नुहुन्छ।

Alibaba का अन्य Models

क्याटलग अद्यावधिक गरिएको मिति: 2026 सेप्टेम्बर 9