99MODELS

मिस्ट्रल स्मॉल 3

Cheap European model with vision input and optional reasoning.

रिलीज़: 16 मार्च 2026

72.00

प्रति 10 लाख आउटपुट Tokens

इनपुट: ₹18.00 प्रति 10 लाख Tokens

प्रदाता की आधिकारिक दर पर बिलिंग, ₹96 प्रति डॉलर पर कनवर्ट, 0% मार्कअप के साथ।

स्पेसिफिकेशन्स

Context विंडो
2,62,144 Tokens
अधिकतम आउटपुट
2,09,715 Tokens
स्वीकार्य इनपुट
टेक्स्ट, इमेज
Reasoning
वैकल्पिक
प्रयास स्तर
none, high
टूल का उपयोग
हाँ
स्ट्रक्चर्ड आउटपुट
हाँ
कोड एग्जीक्यूशन
नहीं
इंटेलिजेंस रैंक
54 में से #51
वैल्यू रैंक
54 में से #47

दरें

दरें
प्रति 10 lakh TokensINRUSD
इनपुट18.00$0.19
आउटपुट72.00$0.75
कैश्ड इनपुट1.58$0.02

Benchmarks

स्कोर प्रतिशत में हैं जब तक कि कोई रेटिंग न दी गई हो। सभी Benchmark स्वतंत्र रूप से मापे गए हैं।

  • 76.9%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 9.9%

    HLE

    Humanity's Last Exam

  • 48.2%

    IFBench

    IFBench - precise instruction following

  • 49.7%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 17.4%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 21.0%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

मिस्ट्रल स्मॉल 3 के बारे में

Mistral Small 4, the lab's hybrid model unifying instruct, reasoning and coding in one efficient set of weights. It is the first Mistral to fold the Magistral reasoning, Pixtral multimodal and Devstral agentic-coding lines into a single versatile model, tuned for general chat, coding, agentic tasks and complex reasoning. It is a Mixture-of-Experts with 119B total and roughly 6B active parameters, taking text and image input across a 256K context.

Small 4 is the release where Mistral stopped asking people to pick a model. The lab describes it as the first Mistral to fold its reasoning, multimodal and agentic-coding lines into one set of weights, so a fast instruct reply, a step-by-step derivation and a tool-driven coding run all come from the same checkpoint. Its sparse design routes four of 128 experts per token, which is why a 119 billion parameter model costs what a much smaller one does to serve.

The control that makes the merge work is a reasoning-effort parameter with two named settings, and Mistral describes both in terms of the models they replace. Setting it to none produces fast, lightweight answers in the same chat style as the previous Small release. Setting it to high produces deep step-by-step reasoning with roughly the verbosity of the lab's dedicated reasoning models. That is the whole switch: no second endpoint, no second deployment.

Mistral's own comparison is against its earlier models rather than the field. With reasoning on it reports 71.2 on GPQA Diamond and 78 on MMLU Pro, against 59.1 and 73.5 for the same weights in instruct mode, 48 against 35.7 on the AllenAI instruction-following benchmark IFBench, and 60 against 46.3 on the vision benchmark MMMU-Pro. Alongside the accuracy the lab pushes an efficiency claim: it says Small 4 matches or beats a comparable open model on three benchmarks while generating substantially shorter answers, and that shorter answers are the point, because they are what actually shows up as latency and cost.

The engineering numbers are stated plainly. Mistral reports a 40% cut in end-to-end completion time in a latency-tuned setup and three times the requests per second in a throughput-tuned one, against the previous Small generation. It names the minimum hardware - four H100s, two H200s or a single DGX B200 - and ships under Apache 2.0, with day-one support across the common open serving stacks. The lab's own framing of who it is for is equally plain: developers doing coding automation and codebase exploration, enterprises running chat assistants and document understanding, researchers doing maths and complex reasoning. It is the efficient tier, not the frontier one, and Mistral points at its larger models for the hardest work.

लॉन्च के समय Mistral AI ने क्या कहा

Three model lines in one
Mistral's first model to unify its reasoning, multimodal and agentic-coding lines, so users no longer choose between a fast instruct model, a reasoning engine and a vision assistant.
Reasoning effort as a parameter
Setting effort to none gives the previous generation chat style; setting it to high gives step-by-step reasoning with the verbosity of the dedicated reasoning models it replaces.
Sparse routing at 119B
A Mixture-of-Experts with 128 experts and four active per token, around 6B active parameters per token, which is what keeps serving cost near a much smaller model.
Reasoning versus instruct scores
The lab reports 71.2 against 59.1 on GPQA Diamond, 78 against 73.5 on MMLU Pro and 60 against 46.3 on MMMU-Pro, comparing the same weights with reasoning on and off.
Shorter answers on purpose
Mistral argues efficiency per token is the real metric and reports matching or beating a comparable open model while generating substantially less text to get there.
Serving cost and hardware floor
The lab reports a 40% reduction in end-to-end completion time and three times the requests per second against the previous Small generation, with a floor of four H100s, two H200s or one DGX B200.

भारतीय भाषाएं

मिस्ट्रल स्मॉल 3 4 भारतीय भाषाओं में जवाब दे सकता है। मैसेज बॉक्स के पास वाले मेन्यू से भाषा चुनें और उसी में जवाब पाएं।

Frequently Asked Questions

मिस्ट्रल स्मॉल 3 के बारे में अक्सर पूछे जाने वाले सवाल।

मिस्ट्रल स्मॉल 3 कब रिलीज़ हुआ था?

Mistral AI ने मिस्ट्रल स्मॉल 3 को 16 मार्च 2026 को रिलीज़ किया था।

मिस्ट्रल स्मॉल 3 को किसने बनाया है?

मिस्ट्रल स्मॉल 3 को AI लैब Mistral AI ने बनाया है। 99Models इसे सीधे प्रदाता की आधिकारिक दर पर उपलब्ध कराता है।

मिस्ट्रल स्मॉल 3 कितना समझदार है?

इंटेलीजेंस रैंकिंग में यह 54 चैट Models में से 51 स्थान पर है, जो बेंचमार्क स्कोर पर आधारित है। इसके पूरे स्कोर ऊपर Benchmarks पैनल में देखे जा सकते हैं।

मिस्ट्रल स्मॉल 3 का उपयोग करने का क्या खर्च है?

10 लाख इनपुट Tokens के लिए ₹18.00 और 10 लाख आउटपुट Tokens के लिए ₹72.00, बिना किसी अतिरिक्त मार्कअप के प्रदाता की दर पर। कोई सब्सक्रिप्शन नहीं है; आप केवल अपने उपयोग का भुगतान करते हैं।

डॉलर में मिस्ट्रल स्मॉल 3 API की कीमत क्या है?

प्रदाता 10 लाख इनपुट Tokens के लिए $0.19 और 10 लाख आउटपुट Tokens के लिए $0.75 चार्ज करता है। इस पेज पर रुपये की दरें ₹96 प्रति डॉलर के हिसाब से बदली गई हैं।

मिस्ट्रल स्मॉल 3 कितनी लंबी बातचीत याद रख सकता है?

इसकी Context विंडो 2.6 lakh Tokens है। यानी यह एक रिक्वेस्ट में पिछली बातचीत और अटैच की गई फ़ाइलों को मिलाकर इतना टेक्स्ट पढ़ सकता है।

क्या मिस्ट्रल स्मॉल 3 पैसे के लिहाज से किफ़ायती है?

वैल्यू रैंकिंग में यह 54 Models में से 47 स्थान पर है। यह रैंकिंग परफ़ॉर्मेंस और टोकन की वास्तविक कीमत की तुलना करके तय की जाती है।

क्या मिस्ट्रल स्मॉल 3 जवाब देने से पहले सोच-विचार (Reasoning) करता है?

वैकल्पिक। जहां सोचने की क्षमता उपलब्ध है, वहां आप मैसेज कंपोज़र में सीधे Reasoning का स्तर तय कर सकते हैं।

मिस्ट्रल स्मॉल 3 किन भारतीय भाषाओं में जवाब दे सकता है?

यह 4 भारतीय भाषाओं में जवाब देता है। मैसेज बॉक्स के पास वाले मेन्यू से भाषा चुनें और उसी भाषा में जवाब पाएं।

Mistral AI के अन्य Models

कैटलॉग अपडेट: 9 सित॰ 2026