99MODELS

Mistral Small 3

Cheap European model with vision input and optional reasoning.

Released Mar 16, 2026

72.00

per 10 lakh output tokens

Input: ₹18.00 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
2,62,144 tokens
Max output
2,09,715 tokens
Accepts
Text, Images
Reasoning
Optional
Effort levels
none, high
Tool use
Yes
Structured output
Yes
Code execution
No
Intelligence rank
#51 of 54
Value rank
#47 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input18.00$0.19
Output72.00$0.75
Cached input1.58$0.02

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 76.9%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 9.9%

    HLE

    Humanity's Last Exam

  • 48.2%

    IFBench

    IFBench - precise instruction following

  • 49.7%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 17.4%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 21.0%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

About Mistral Small 3

Mistral Small 4, the lab's hybrid model unifying instruct, reasoning and coding in one efficient set of weights. It is the first Mistral to fold the Magistral reasoning, Pixtral multimodal and Devstral agentic-coding lines into a single versatile model, tuned for general chat, coding, agentic tasks and complex reasoning. It is a Mixture-of-Experts with 119B total and roughly 6B active parameters, taking text and image input across a 256K context.

Small 4 is the release where Mistral stopped asking people to pick a model. The lab describes it as the first Mistral to fold its reasoning, multimodal and agentic-coding lines into one set of weights, so a fast instruct reply, a step-by-step derivation and a tool-driven coding run all come from the same checkpoint. Its sparse design routes four of 128 experts per token, which is why a 119 billion parameter model costs what a much smaller one does to serve.

The control that makes the merge work is a reasoning-effort parameter with two named settings, and Mistral describes both in terms of the models they replace. Setting it to none produces fast, lightweight answers in the same chat style as the previous Small release. Setting it to high produces deep step-by-step reasoning with roughly the verbosity of the lab's dedicated reasoning models. That is the whole switch: no second endpoint, no second deployment.

Mistral's own comparison is against its earlier models rather than the field. With reasoning on it reports 71.2 on GPQA Diamond and 78 on MMLU Pro, against 59.1 and 73.5 for the same weights in instruct mode, 48 against 35.7 on the AllenAI instruction-following benchmark IFBench, and 60 against 46.3 on the vision benchmark MMMU-Pro. Alongside the accuracy the lab pushes an efficiency claim: it says Small 4 matches or beats a comparable open model on three benchmarks while generating substantially shorter answers, and that shorter answers are the point, because they are what actually shows up as latency and cost.

The engineering numbers are stated plainly. Mistral reports a 40% cut in end-to-end completion time in a latency-tuned setup and three times the requests per second in a throughput-tuned one, against the previous Small generation. It names the minimum hardware - four H100s, two H200s or a single DGX B200 - and ships under Apache 2.0, with day-one support across the common open serving stacks. The lab's own framing of who it is for is equally plain: developers doing coding automation and codebase exploration, enterprises running chat assistants and document understanding, researchers doing maths and complex reasoning. It is the efficient tier, not the frontier one, and Mistral points at its larger models for the hardest work.

What Mistral AI announced at launch

Three model lines in one
Mistral's first model to unify its reasoning, multimodal and agentic-coding lines, so users no longer choose between a fast instruct model, a reasoning engine and a vision assistant.
Reasoning effort as a parameter
Setting effort to none gives the previous generation chat style; setting it to high gives step-by-step reasoning with the verbosity of the dedicated reasoning models it replaces.
Sparse routing at 119B
A Mixture-of-Experts with 128 experts and four active per token, around 6B active parameters per token, which is what keeps serving cost near a much smaller model.
Reasoning versus instruct scores
The lab reports 71.2 against 59.1 on GPQA Diamond, 78 against 73.5 on MMLU Pro and 60 against 46.3 on MMMU-Pro, comparing the same weights with reasoning on and off.
Shorter answers on purpose
Mistral argues efficiency per token is the real metric and reports matching or beating a comparable open model while generating substantially less text to get there.
Serving cost and hardware floor
The lab reports a 40% reduction in end-to-end completion time and three times the requests per second against the previous Small generation, with a floor of four H100s, two H200s or one DGX B200.

Indian languages

Mistral Small 3 answers in 4 Indian languages. Choose a language from the menu beside the message box to receive replies in it.

Frequently Asked Questions

Frequently asked questions about Mistral Small 3.

When was Mistral Small 3 released?

Mistral AI released Mistral Small 3 on Mar 16, 2026.

Who built Mistral Small 3?

Mistral Small 3 is developed by Mistral AI. 99Models AI connects directly to it at the provider's published rate.

How intelligent is Mistral Small 3?

It is ranked 51 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does Mistral Small 3 cost?

Usage costs ₹18.00 per million input tokens and ₹72.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is Mistral Small 3 pricing in US dollars?

The provider charges $0.19 per million input tokens and $0.75 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can Mistral Small 3 hold?

Its context window is 2.6 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does Mistral Small 3 rank for value?

It ranks 47 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does Mistral Small 3 support reasoning?

Optional. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

Which Indian languages does Mistral Small 3 support?

It answers in 4 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.

More models from Mistral AI

Catalog updated Sep 9, 2026