Qwen 3.8 Max
NewAlibaba-hosted 3.8 flagship; 256K window on our one route.
Released Sep 3, 2026
₹576.00
per 10 lakh output tokens
Input: ₹192.00 per 10 lakh tokens
Billed at provider rates converted at ₹96 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
92.7%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
43.0%
HLE
Humanity's Last Exam
53.2%
SciCode
SciCode - scientific code generation
78.3%
Long Context
Long Context Reasoning - reasoning over long inputs
81.3%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
About Qwen 3.8 Max
Alibaba's proprietary 3.8 flagship: a 2.4-trillion-parameter sparse Mixture-of-Experts model positioned as a comprehensive step up in coding and professional knowledge work over the previous Max generation. It is the hosted, closed counterpart to the openly released 3.8 2.4T, and it always reasons. Note that the window here is 256K rather than the million tokens the model is capable of, because the one host we route to serves it at that length.
Alibaba argued this model with long autonomous runs rather than with single scores, and the three coding cases in its launch post are all days long. Asked to create a command-line project from an empty folder, the model ran for more than ten days building a harness that folds community feedback, its own test results and normalised issues into one loop; after roughly sixteen days of unattended operation the lab counted 265 commits, 127 pull requests and 151 issues in the repository. Given a research paper, a set of GPUs and no starter code, it worked about 125 hours, wrote roughly 7,600 lines and ran 33 rounds of GPU training, reproducing the paper's six findings in the first 37 hours and then inventing and testing 18 ideas of its own across four rounds until it beat the paper's own method by 2.7 points on AIME24. Entered into a live online contest against 526 human teams under a 24-hour limit, it climbed from 0.60 to 0.853 accuracy over 45 submissions and finished ahead of 458 of them.
The thread Alibaba draws through all three is that the model does not just execute a plan - it revises the plan from feedback. That is also how the lab describes its training for ordinary work: reinforcement-learning environments and compute scaled together along task, workspace and harness axes, with one reward system spanning execution checks, rubric judging of both text and rendered visual output, and agentic inspection, plus a batch balancer to keep training stable as it scales. The stated aim is competence that lifts across harnesses rather than in one, and Alibaba reports comparable results across several popular agent harnesses.
It then stress-tested breadth across several hundred high-value professions. The examples the lab publishes are concrete: a compliance review that surfaced 1,284 relevant clauses across hundreds of documents in under an hour, an eight-screen banking prototype with a consistent design system delivered in one shot with no revision rounds, a 26-dish menu costed to a 33.8% food-cost ratio from a hundred-odd supply briefs, and a seismic model of a 30-storey tower reconstructed in the browser from one set of drawings.
For long-horizon work Alibaba reports two results worth quoting. On a 365-day e-commerce operations simulation built from real transaction data, with nearly 600 suppliers of which 152 were deliberately fraudulent, the model ended the year with the highest balance of the field, 416,252 yuan from 100,000 in starting capital - 38% ahead of the next model and 152% above the previous Max generation. On an autonomous chip-design task it took a cryptographic accelerator from 8,298 gates to 678 across roughly 500 turns, with the physical die shrinking 81% after place and route, and Alibaba notes the largest single structural improvement came hundreds of turns into the run rather than early.
On its own benchmark table Alibaba reports the best score in the comparison on PaperBench at 93.0, WideSearch at 81.9, IFBench at 82.8 and HealthBench at 60.2, alongside 86.6 on Terminal Bench 2.1, 92.6 on GPQA Diamond and 43.6 on HLE. It also publishes where it trails: 56.6 on DeepSWE v1.1 against 73.0 and 70.0 for the two models it compares with, and 67.7 on SWE-bench Pro against 80.0. Reasoning is always on and the effort dial defaults to its highest setting, with reasoning from earlier turns preserved by default; the lab's own caution is that on multi-turn agentic work a lower effort setting does not reliably finish sooner, because thinner analysis produces more failures and retries.
What Alibaba announced at launch
- Ten days unattended
- Alibaba reports a single autonomous run of more than ten days that produced 265 commits, 127 pull requests and 151 issues while the model maintained and extended its own harness.
- Reproducing and improving a paper
- From nothing but a paper and GPUs, the lab says it rebuilt the full pipeline in about 37 hours, then tested 18 ideas of its own and beat the paper's method by 2.7 points on AIME24.
- Work across professions
- Alibaba publishes examples from several hundred high-value professions, including 1,284 contract clauses surfaced in under an hour and an eight-screen prototype delivered with no revision rounds.
- Long-horizon operations
- On a 365-day business simulation the lab reports the highest final balance in the field, 38% ahead of the next model, and on a chip-design task a reduction from 8,298 gates to 678.
- Benchmarks the lab leads
- Alibaba reports the best score in its comparison on PaperBench at 93.0, WideSearch at 81.9, IFBench at 82.8 and HealthBench at 60.2, with 92.6 on GPQA Diamond.
- What it is not for
- The same table has it behind on DeepSWE v1.1 at 56.6 against 73.0, and on SWE-bench Pro at 67.7 against 80.0. Reasoning cannot be disabled, so there is no fast non-thinking mode.
Frequently Asked Questions
Frequently asked questions about Qwen 3.8 Max.
When was Qwen 3.8 Max released?
Alibaba released Qwen 3.8 Max on Sep 3, 2026.
Who built Qwen 3.8 Max?
Qwen 3.8 Max is developed by Alibaba. 99Models AI connects directly to it at the provider's published rate.
How intelligent is Qwen 3.8 Max?
It is ranked 13 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does Qwen 3.8 Max cost?
Usage costs ₹192.00 per million input tokens and ₹576.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is Qwen 3.8 Max pricing in US dollars?
The provider charges $2.00 per million input tokens and $6.00 per million output tokens. Rupee rates are converted at ₹96 per US dollar.
How long a conversation can Qwen 3.8 Max hold?
Its context window is 2.6 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does Qwen 3.8 Max rank for value?
It ranks 14 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does Qwen 3.8 Max support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
More models from Alibaba
Catalog updated Sep 9, 2026