99MODELS

Qwen 3.8 Max

New

Alibaba-hosted 3.8 flagship; 256K window on our one route.

Released Sep 3, 2026

576.00

per 10 lakh output tokens

Input: ₹192.00 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
2,56,000 tokens
Max output
1,31,072 tokens
Accepts
Text
Reasoning
On by default
Effort levels
minimal, low, medium, high, xhigh
Tool use
Yes
Structured output
Yes
Code execution
No
Parameters
2.4T MoE
Intelligence rank
#13 of 54
Value rank
#14 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input192.00$2.00
Output576.00$6.00
Cached input24.00$0.25

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 92.7%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 43.0%

    HLE

    Humanity's Last Exam

  • 53.2%

    SciCode

    SciCode - scientific code generation

  • 78.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 81.3%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

About Qwen 3.8 Max

Alibaba's proprietary 3.8 flagship: a 2.4-trillion-parameter sparse Mixture-of-Experts model positioned as a comprehensive step up in coding and professional knowledge work over the previous Max generation. It is the hosted, closed counterpart to the openly released 3.8 2.4T, and it always reasons. Note that the window here is 256K rather than the million tokens the model is capable of, because the one host we route to serves it at that length.

Alibaba argued this model with long autonomous runs rather than with single scores, and the three coding cases in its launch post are all days long. Asked to create a command-line project from an empty folder, the model ran for more than ten days building a harness that folds community feedback, its own test results and normalised issues into one loop; after roughly sixteen days of unattended operation the lab counted 265 commits, 127 pull requests and 151 issues in the repository. Given a research paper, a set of GPUs and no starter code, it worked about 125 hours, wrote roughly 7,600 lines and ran 33 rounds of GPU training, reproducing the paper's six findings in the first 37 hours and then inventing and testing 18 ideas of its own across four rounds until it beat the paper's own method by 2.7 points on AIME24. Entered into a live online contest against 526 human teams under a 24-hour limit, it climbed from 0.60 to 0.853 accuracy over 45 submissions and finished ahead of 458 of them.

The thread Alibaba draws through all three is that the model does not just execute a plan - it revises the plan from feedback. That is also how the lab describes its training for ordinary work: reinforcement-learning environments and compute scaled together along task, workspace and harness axes, with one reward system spanning execution checks, rubric judging of both text and rendered visual output, and agentic inspection, plus a batch balancer to keep training stable as it scales. The stated aim is competence that lifts across harnesses rather than in one, and Alibaba reports comparable results across several popular agent harnesses.

It then stress-tested breadth across several hundred high-value professions. The examples the lab publishes are concrete: a compliance review that surfaced 1,284 relevant clauses across hundreds of documents in under an hour, an eight-screen banking prototype with a consistent design system delivered in one shot with no revision rounds, a 26-dish menu costed to a 33.8% food-cost ratio from a hundred-odd supply briefs, and a seismic model of a 30-storey tower reconstructed in the browser from one set of drawings.

For long-horizon work Alibaba reports two results worth quoting. On a 365-day e-commerce operations simulation built from real transaction data, with nearly 600 suppliers of which 152 were deliberately fraudulent, the model ended the year with the highest balance of the field, 416,252 yuan from 100,000 in starting capital - 38% ahead of the next model and 152% above the previous Max generation. On an autonomous chip-design task it took a cryptographic accelerator from 8,298 gates to 678 across roughly 500 turns, with the physical die shrinking 81% after place and route, and Alibaba notes the largest single structural improvement came hundreds of turns into the run rather than early.

On its own benchmark table Alibaba reports the best score in the comparison on PaperBench at 93.0, WideSearch at 81.9, IFBench at 82.8 and HealthBench at 60.2, alongside 86.6 on Terminal Bench 2.1, 92.6 on GPQA Diamond and 43.6 on HLE. It also publishes where it trails: 56.6 on DeepSWE v1.1 against 73.0 and 70.0 for the two models it compares with, and 67.7 on SWE-bench Pro against 80.0. Reasoning is always on and the effort dial defaults to its highest setting, with reasoning from earlier turns preserved by default; the lab's own caution is that on multi-turn agentic work a lower effort setting does not reliably finish sooner, because thinner analysis produces more failures and retries.

What Alibaba announced at launch

Ten days unattended
Alibaba reports a single autonomous run of more than ten days that produced 265 commits, 127 pull requests and 151 issues while the model maintained and extended its own harness.
Reproducing and improving a paper
From nothing but a paper and GPUs, the lab says it rebuilt the full pipeline in about 37 hours, then tested 18 ideas of its own and beat the paper's method by 2.7 points on AIME24.
Work across professions
Alibaba publishes examples from several hundred high-value professions, including 1,284 contract clauses surfaced in under an hour and an eight-screen prototype delivered with no revision rounds.
Long-horizon operations
On a 365-day business simulation the lab reports the highest final balance in the field, 38% ahead of the next model, and on a chip-design task a reduction from 8,298 gates to 678.
Benchmarks the lab leads
Alibaba reports the best score in its comparison on PaperBench at 93.0, WideSearch at 81.9, IFBench at 82.8 and HealthBench at 60.2, with 92.6 on GPQA Diamond.
What it is not for
The same table has it behind on DeepSWE v1.1 at 56.6 against 73.0, and on SWE-bench Pro at 67.7 against 80.0. Reasoning cannot be disabled, so there is no fast non-thinking mode.

Frequently Asked Questions

Frequently asked questions about Qwen 3.8 Max.

When was Qwen 3.8 Max released?

Alibaba released Qwen 3.8 Max on Sep 3, 2026.

Who built Qwen 3.8 Max?

Qwen 3.8 Max is developed by Alibaba. 99Models AI connects directly to it at the provider's published rate.

How intelligent is Qwen 3.8 Max?

It is ranked 13 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does Qwen 3.8 Max cost?

Usage costs ₹192.00 per million input tokens and ₹576.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is Qwen 3.8 Max pricing in US dollars?

The provider charges $2.00 per million input tokens and $6.00 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can Qwen 3.8 Max hold?

Its context window is 2.6 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does Qwen 3.8 Max rank for value?

It ranks 14 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does Qwen 3.8 Max support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

More models from Alibaba

Catalog updated Sep 9, 2026