99MODELS

Qwen 3.8 27B

New

Dense 27B Qwen with vision and a 1M-token context.

Released Aug 14, 2026

307.20

per 10 lakh output tokens

Input: ₹43.20 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
10,00,000 tokens
Max output
32,768 tokens
Accepts
Text, Images, Video
Reasoning
On by default
Effort levels
low, medium, xhigh
Tool use
Yes
Structured output
Yes
Code execution
No
Parameters
27B
Intelligence rank
#25 of 54
Value rank
#12 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input43.20$0.45
Output307.20$3.20
Cached input14.40$0.15

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 90.5%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 33.9%

    HLE

    Humanity's Last Exam

  • 46.6%

    SciCode

    SciCode - scientific code generation

  • 82.0%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 79.8%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 1595

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

About Qwen 3.8 27B

Qwen's native vision-language dense model, building on the previous 27B generation with key improvements in coding and office productivity. It is a dense 27B model with a built-in vision encoder giving native image and video understanding alongside text, and it is the compact, deployment-friendly member of the 3.8 family. Thinking is on by default and can be disabled per request, with reasoning preserved across conversation history.

Alibaba shipped this one quietly. There is no launch post for the 27B: it was announced on 14 August 2026 through the lab's own Qwen3.8 repository and its model card, two days after the Max-class open weights, as the compact member of the same generation. Everything below therefore comes from the lab's own release notes and model card rather than from a blog write-up, and there is no marketing framing to discount.

It is a dense model, and that is the point of it. All 27 billion parameters run on every token, which costs more per token than a sparse model of similar size but makes it far simpler to deploy - one machine, predictable memory, no expert routing to tune. Alibaba describes 64 layers in a hybrid layout that runs three Gated DeltaNet blocks for every one full-attention block, with multi-token prediction trained in, native context of 262,144 tokens and extension to a million.

Vision is native rather than attached. The 27B is a vision-language model with its own encoder, handling images and video from STEM diagrams and scanned documents through to hour-scale footage, and the lab's strongest reported results are on that side of the ledger: 84.3 on OSWorld-Verified for computer use, 81.9 on AndroidWorld for phone use, 64.8 on WebArena-Verified for browser use, 90.0 on MathVision and 91.1 on document parsing. On several of these Alibaba reports it ahead of the much larger previous-generation model it replaces.

Thinking is on by default and can be switched off per request, the effort dial goes from low up to the highest setting, and reasoning from earlier turns is kept in context by default - which Alibaba says both keeps agent decisions consistent and improves cache reuse. The lab adds a caution worth heeding: on multi-turn agentic tasks a lower effort setting does not reliably finish faster, because thinner analysis leads to more failures and retries.

Where it stops is exactly where you would expect the compact member to stop. Alibaba's own text table puts it at 89.2 on GPQA Diamond and 30.8 on HLE, against 92.6 and 43.6 for the Max-class model in the same family, so the ceiling on hard general knowledge is materially lower. Its coding results - 73.0 on Terminal Bench 2.1 and 61.7 on SWE-bench Pro - clear the previous 27B and the previous flagship comfortably but sit below the frontier. The compensation is that it is released under a permissive Apache-2.0 licence, unlike its Max-class sibling.

What Alibaba announced at launch

Dense, not sparse
All 27 billion parameters run on every token, which is heavier per token than a sparse model but far simpler to deploy on a single machine with predictable memory.
Vision is native
It carries its own vision encoder and handles images and video, from STEM diagrams and scanned documents through to hour-scale footage, rather than bolting understanding on.
Computer, phone and browser use
Alibaba reports 84.3 on OSWorld-Verified, 81.9 on AndroidWorld and 64.8 on WebArena-Verified, ahead of the much larger previous-generation model on several of them.
Thinking on by default
It thinks unless told not to, keeps reasoning from earlier turns in context, and exposes an effort dial - though the lab warns lower effort does not reliably finish agentic tasks sooner.
Apache-2.0
Unlike the Max-class checkpoint in the same family, this one is released under permissive terms, which matters if you intend to build a product on it.
What it is not for
Alibaba's own table puts it at 89.2 on GPQA Diamond and 30.8 on HLE against 92.6 and 43.6 for the Max-class model, so hard general reasoning belongs somewhere else.

Indian languages

Qwen 3.8 27B answers in 1 Indian languages. Choose a language from the menu beside the message box to receive replies in it.

Frequently Asked Questions

Frequently asked questions about Qwen 3.8 27B.

When was Qwen 3.8 27B released?

Alibaba released Qwen 3.8 27B on Aug 14, 2026.

Who built Qwen 3.8 27B?

Qwen 3.8 27B is developed by Alibaba. 99Models AI connects directly to it at the provider's published rate.

How intelligent is Qwen 3.8 27B?

It is ranked 25 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does Qwen 3.8 27B cost?

Usage costs ₹43.20 per million input tokens and ₹307.20 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is Qwen 3.8 27B pricing in US dollars?

The provider charges $0.45 per million input tokens and $3.20 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can Qwen 3.8 27B hold?

Its context window is 10 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does Qwen 3.8 27B rank for value?

It ranks 12 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does Qwen 3.8 27B support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

Which Indian languages does Qwen 3.8 27B support?

It answers in 1 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.

More models from Alibaba

Catalog updated Sep 9, 2026