क्वेन 3.8 2.4T
नयाOpen-weight 2.4T MoE flagship; near-frontier reasoning.
रिलीज़: 12 अग॰ 2026
₹720.00
प्रति 10 लाख आउटपुट Tokens
इनपुट: ₹240.00 प्रति 10 लाख Tokens
प्रदाता की आधिकारिक दर पर बिलिंग, ₹96 प्रति डॉलर पर कनवर्ट, 0% मार्कअप के साथ।
स्पेसिफिकेशन्स
- Context विंडो
- 10,48,576 Tokens
- अधिकतम आउटपुट
- 9,09,000 Tokens
- स्वीकार्य इनपुट
- टेक्स्ट
- Reasoning
- डिफ़ॉल्ट रूप से चालू
- प्रयास स्तर
- low, medium, xhigh
- टूल का उपयोग
- हाँ
- स्ट्रक्चर्ड आउटपुट
- हाँ
- कोड एग्जीक्यूशन
- नहीं
- पैरामीटर
- 2.4T A95B MoE
- इंटेलिजेंस रैंक
- 54 में से #16
- वैल्यू रैंक
- 54 में से #18
दरें
Benchmarks
स्कोर प्रतिशत में हैं जब तक कि कोई रेटिंग न दी गई हो। सभी Benchmark स्वतंत्र रूप से मापे गए हैं।
93.5%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
42.4%
HLE
Humanity's Last Exam
54.1%
SciCode
SciCode - scientific code generation
80.3%
Long Context
Long Context Reasoning - reasoning over long inputs
82.0%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
क्वेन 3.8 2.4T के बारे में
The open-source release of Qwen's latest flagship, and the first time a Qwen Max-class model has been published openly. Its sparse Mixture-of-Experts architecture holds 2.4T total parameters with roughly 95B activated per step, combining Gated DeltaNet and Gated Attention across 92 layers with a million-token context. It always reasons: every response begins with an internal reasoning pass before the final output.
This is the checkpoint the Qwen3.8-Max announcement promised. Alibaba used that post to say it would open-source the weights of a Max-class model for the first time, and the weights landed nine days later, on 12 August 2026, with no separate write-up. For a reader the practical meaning is straightforward: this is the same post-trained model behind the lab's hosted flagship, downloadable and runnable on hardware you choose rather than reachable only through one vendor's API. The licence is the caveat - it ships under a custom Qwen3.8-Max licence rather than the permissive terms Qwen's smaller open models carry, so commercial terms are worth reading before building on it.
The architecture is what makes a model this size servable at all. Alibaba describes 92 layers holding 2.4 trillion parameters with roughly 95 billion active per token, drawn from 512 experts of which ten routed plus one shared fire on any given token, and a hybrid backbone that runs Gated DeltaNet linear attention in three of every four layers with a full Gated Attention layer in the fourth. Multi-token prediction is trained in. Context is 262,144 tokens natively and extends past a million.
Alibaba is unusually explicit about what this checkpoint does not do, and the constraints are real rather than cosmetic. It is text only: multimodal input is not supported. Thinking cannot be disabled - every response begins with an internal reasoning pass before the final output, at any effort setting. The hosted sibling is the version that adds image input, a non-thinking mode and the million-token window by default, so the open weights are the reasoning core of the product rather than the whole product.
Because it is the same model, its reported numbers are the flagship's: Alibaba publishes 86.6 on Terminal Bench 2.1, 67.7 on SWE-bench Pro, 93.0 on PaperBench, 92.6 on GPQA Diamond, 43.6 on HLE, 82.8 on IFBench and 92.9 on the 256K eight-needle retrieval test. The same table shows it behind on DeepSWE v1.1 at 56.6 against 73.0 and 70.0, and on SWE-bench Pro against a leading 80.0.
Running it well takes some budgeting. The effort dial defaults to its highest setting, thinking from earlier turns is preserved by default - which Alibaba says helps agent consistency and cache reuse - and the lab's own recommendation inside a million-token context is to allow 262,144 tokens for reasoning and 131,072 for the final response. That split is the clearest statement of what this model is: one that spends a large share of its context thinking before it writes anything down.
लॉन्च के समय Alibaba ने क्या कहा
- First open Max-class model
- Alibaba had never published the weights of a Max-class Qwen before this release, which is the claim the launch post leads with.
- Sparse by design
- 2.4 trillion total parameters with about 95 billion active per token, from 512 experts firing ten routed and one shared, across a 92-layer hybrid attention backbone.
- Text only, always thinking
- Alibaba states plainly that multimodal input is not supported and thinking cannot be disabled - every response starts with an internal reasoning pass.
- The licence is not permissive
- It ships under a custom Qwen3.8-Max licence rather than the permissive terms the lab's smaller open models use, so commercial use needs reading before it is assumed.
- Budget the context for thinking
- Inside a million-token window the lab recommends allowing 262,144 tokens for reasoning and 131,072 for the final answer, which says a lot about how it spends a request.
- Where it trails
- On Alibaba's own table it sits at 56.6 on DeepSWE v1.1 against 73.0 and 70.0 for the models it compares with, and at 67.7 on SWE-bench Pro against a leading 80.0.
भारतीय भाषाएं
क्वेन 3.8 2.4T 12 भारतीय भाषाओं में जवाब दे सकता है। मैसेज बॉक्स के पास वाले मेन्यू से भाषा चुनें और उसी में जवाब पाएं।
Frequently Asked Questions
क्वेन 3.8 2.4T के बारे में अक्सर पूछे जाने वाले सवाल।
क्वेन 3.8 2.4T कब रिलीज़ हुआ था?
Alibaba ने क्वेन 3.8 2.4T को 12 अग॰ 2026 को रिलीज़ किया था।
क्वेन 3.8 2.4T को किसने बनाया है?
क्वेन 3.8 2.4T को AI लैब Alibaba ने बनाया है। 99Models इसे सीधे प्रदाता की आधिकारिक दर पर उपलब्ध कराता है।
क्वेन 3.8 2.4T कितना समझदार है?
इंटेलीजेंस रैंकिंग में यह 54 चैट Models में से 16 स्थान पर है, जो बेंचमार्क स्कोर पर आधारित है। इसके पूरे स्कोर ऊपर Benchmarks पैनल में देखे जा सकते हैं।
क्वेन 3.8 2.4T का उपयोग करने का क्या खर्च है?
10 लाख इनपुट Tokens के लिए ₹240.00 और 10 लाख आउटपुट Tokens के लिए ₹720.00, बिना किसी अतिरिक्त मार्कअप के प्रदाता की दर पर। कोई सब्सक्रिप्शन नहीं है; आप केवल अपने उपयोग का भुगतान करते हैं।
डॉलर में क्वेन 3.8 2.4T API की कीमत क्या है?
प्रदाता 10 लाख इनपुट Tokens के लिए $2.50 और 10 लाख आउटपुट Tokens के लिए $7.50 चार्ज करता है। इस पेज पर रुपये की दरें ₹96 प्रति डॉलर के हिसाब से बदली गई हैं।
क्वेन 3.8 2.4T कितनी लंबी बातचीत याद रख सकता है?
इसकी Context विंडो 10.5 lakh Tokens है। यानी यह एक रिक्वेस्ट में पिछली बातचीत और अटैच की गई फ़ाइलों को मिलाकर इतना टेक्स्ट पढ़ सकता है।
क्या क्वेन 3.8 2.4T पैसे के लिहाज से किफ़ायती है?
वैल्यू रैंकिंग में यह 54 Models में से 18 स्थान पर है। यह रैंकिंग परफ़ॉर्मेंस और टोकन की वास्तविक कीमत की तुलना करके तय की जाती है।
क्या क्वेन 3.8 2.4T जवाब देने से पहले सोच-विचार (Reasoning) करता है?
डिफ़ॉल्ट रूप से चालू। जहां सोचने की क्षमता उपलब्ध है, वहां आप मैसेज कंपोज़र में सीधे Reasoning का स्तर तय कर सकते हैं।
क्वेन 3.8 2.4T किन भारतीय भाषाओं में जवाब दे सकता है?
यह 12 भारतीय भाषाओं में जवाब देता है। मैसेज बॉक्स के पास वाले मेन्यू से भाषा चुनें और उसी भाषा में जवाब पाएं।
Alibaba के अन्य Models
कैटलॉग अपडेट: 9 सित॰ 2026