99MODELS

Ministral 8B

Tiny edge-class Mistral for classification and routing.

Released Dec 2, 2025

15.84

per 10 lakh output tokens

Input: ₹15.84 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
2,62,144 tokens
Max output
2,09,715 tokens
Accepts
Text, Images
Reasoning
No
Tool use
Yes
Structured output
Yes
Code execution
No
Parameters
8B
Intelligence rank
#54 of 54
Value rank
#50 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input15.84$0.17
Output15.84$0.17
Cached input1.58$0.02

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 64.2%

    MMLU-Pro

    MMLU-Pro - multitask language understanding

  • 47.1%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 4.3%

    HLE

    Humanity's Last Exam

  • 30.3%

    LiveCodeBench

    LiveCodeBench - contamination-free coding

  • 31.7%

    AIME 2025

    AIME 2025 - competition mathematics

  • 29.1%

    IFBench

    IFBench - precise instruction following

  • 25.7%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 4.5%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 4.1%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

About Ministral 8B

A powerful and efficient model in Mistral's Ministral 3 family, offering best-in-class text and vision capability at its size. It is built for edge deployment and performs across diverse hardware, fitting in 12GB of VRAM at FP8 and less when further quantised, while carrying a 256K context and image understanding across dozens of languages. It does not reason -- the reasoning capability ships as a separate checkpoint rather than a runtime toggle.

The 8B sits in the middle of a three-size edge family - 3B, 8B and 14B - released as the small end of the Mistral 3 generation. Mistral gave every size three separate checkpoints rather than one switchable model: a base, an instruct and a reasoning variant, each with image understanding, all under Apache 2.0. That is the structural fact worth knowing before you pick this one, because it means the reasoning capability is a different download rather than a runtime flag.

Mistral's claim for the family is a ratio rather than a peak: it says the Ministral models offer the best cost-to-performance of any open model, and it argues the case in a way most small-model launches do not. Its point is that in real deployments the tokens generated matter as much as the parameter count, and it reports the instruct variants matching or beating comparable models while often emitting an order of magnitude fewer tokens to get there. The chart it published makes that concrete for this size: the 8B instruct model lands around 51 on GPQA Diamond at roughly 1,500 output tokens, in the same accuracy band as a comparable 8B model that spends more than ten times as many.

The lab is explicit about the division of labour inside the family. For settings where accuracy is the only concern, it points at the reasoning variants, which think longer to reach what it calls state-of-the-art accuracy for their weight class - it cites 85% on AIME 2025 for the 14B. This instruct checkpoint is the opposite trade: answer fast, answer short, stay cheap.

The deployment story is where the size pays. Mistral says the whole Mistral 3 family was trained on NVIDIA Hopper hardware and that it worked with NVIDIA on optimised deployments for desktop AI machines, RTX PCs and laptops, and embedded Jetson boards - the point being that the 8B is meant to run on the device rather than in a datacentre. It carries native multilingual coverage across more than forty languages and image understanding at the same size. What it is not is a frontier model: Mistral positions it for edge inference, classification, routing and on-device assistants, and points at its far larger models for anything that needs frontier reasoning.

What Mistral AI announced at launch

Three sizes, three variants each
Mistral released 3B, 8B and 14B models, and for every size a base, an instruct and a reasoning checkpoint, all with image understanding and all under Apache 2.0.
Accuracy per token, not per parameter
The lab argues generated tokens matter as much as model size in production and reports the instruct models matching comparable ones while often emitting an order of magnitude fewer tokens.
Reasoning is a separate checkpoint
For accuracy-first work Mistral points at the reasoning variants rather than a runtime toggle, citing 85% on AIME 2025 for the 14B member of the family.
Built to run on the device
Mistral worked with NVIDIA on optimised deployments for desktop AI machines, RTX PCs and laptops, and Jetson boards, so the model runs locally rather than in a datacentre.
Multilingual and multimodal at 8B
Native coverage of more than forty languages plus image understanding, both at a size that fits on consumer hardware.
What it is not for
This is the edge tier: Mistral positions the Ministral line for local and cost-sensitive work and points at its far larger models for frontier reasoning.

Frequently Asked Questions

Frequently asked questions about Ministral 8B.

When was Ministral 8B released?

Mistral AI released Ministral 8B on Dec 2, 2025.

Who built Ministral 8B?

Ministral 8B is developed by Mistral AI. 99Models AI connects directly to it at the provider's published rate.

How intelligent is Ministral 8B?

It is ranked 54 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does Ministral 8B cost?

Usage costs ₹15.84 per million input tokens and ₹15.84 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is Ministral 8B pricing in US dollars?

The provider charges $0.17 per million input tokens and $0.17 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can Ministral 8B hold?

Its context window is 2.6 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does Ministral 8B rank for value?

It ranks 50 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

More models from Mistral AI

Catalog updated Sep 9, 2026