99MODELS

MiniMax M3

New

Multimodal 1M-context foundation model for agents and coding.

Released May 31, 2026

230.40

per 10 lakh output tokens

Input: ₹57.60 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
10,48,576 tokens
Max output
4,71,859 tokens
Accepts
Text, Images, Video
Reasoning
On by default
Tool use
Yes
Structured output
Yes
Code execution
No
Intelligence rank
#33 of 54
Value rank
#19 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input57.60$0.60
Output230.40$2.40
Cached input5.76$0.06

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 92.9%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 39.0%

    HLE

    Humanity's Last Exam

  • 47.1%

    SciCode

    SciCode - scientific code generation

  • 82.9%

    IFBench

    IFBench - precise instruction following

  • 83.0%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 42.4%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 65.2%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 14.7%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 22.3%

    OSWorld 2

    OSWorld 2 - agentic computer use (partial credit)

  • 1488

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

About MiniMax M3

MiniMax's native multimodal model with a million-token context, built with roughly 428B total and 23B activated parameters. It uses MiniMax Sparse Attention for efficient long context and is trained on mixed modalities from inception, so text, images and video are integrated semantically rather than having vision bolted on afterwards. MiniMax presents it as the first open-weight model with frontier coding, agentic reasoning and native multimodality at once.

MiniMax makes a combination claim rather than a peak-score one. Frontier coding, a million tokens of context and native multimodality are each unremarkable in a closed frontier model, and M3 is the lab's argument that one set of open weights can carry all three at the same time. What makes that affordable is MSA, the sparse attention design the team proposes: it partitions the key-value cache into blocks more precisely than the alternatives it names, and its operator uses the blocks as the outer loop so each is read exactly once with contiguous memory access.

The efficiency figures are what turn the context from a number into something usable. MiniMax reports the operator running more than four times faster than two open sparse-attention implementations, per-token compute at a million tokens of one twentieth of the previous generation, and end-to-end speedups of more than nine times in prefill and more than fifteen times in decode. Across several ablations it says MSA matched full attention on the vast majority of capabilities, which is the claim that matters most and the easiest one to get wrong.

For coding the lab reports 59.0% on SWE-bench Pro, 66.0% on Terminal-Bench 2.1, 34.8% on SWE-fficiency, 28.8% on the hard tier of a GPU kernel benchmark and 74.2% on MCP Atlas. More interesting than the numbers is the training argument underneath them. MiniMax says most code-agent training and evaluation assumes a single-turn task, which is not how anyone actually works, so it built an interactive user simulator that reproduces requirement elaboration, mid-task correction, task switching and multi-round iteration, and trained and evaluated against that instead.

Its long-run examples follow the same shape. Handed an award-winning machine learning paper and asked to reproduce it independently, M3 ran for close to twelve hours, produced 18 commits and 23 experimental figures and completed the core experiments, including reproducing the effect the paper is known for and verifying the mitigation it proposes. On an FP8 matrix-multiplication kernel with no reference implementation available to imitate, it ran for roughly 24 hours across 147 benchmark submissions and 1,959 tool calls, lifting hardware peak utilisation from 7.6% to 71.3%. MiniMax notes that most other models it tried stopped making progress within the first 30 submissions, while M3's best result arrived on submission 145, after several plateaus it worked through rather than gave up on.

The lab is explicit about where it is not yet ahead. On the post-training benchmark, where the model must autonomously synthesise data, train, evaluate and iterate on four base models inside twelve hours, it scored 0.37, behind Claude Opus 4.7 at 0.42 and GPT-5.5 at 0.39 though clearly ahead of everything else it tested. It describes its agentic performance in the financial domain as only beginning to be usable rather than solved. Weights and a technical report were promised within ten days of the post, so this is a model you could host yourself; the copy served here is MiniMax's.

What MiniMax announced at launch

Three capabilities at once
MiniMax presents M3 as the first open-weight model to combine frontier coding, a million-token context and native multimodality, rather than leading on any one of them alone.
Sparse attention that scales
The lab measures its MSA operator more than four times faster than two open sparse-attention implementations, with prefill over nine times and decode over fifteen times faster at long context.
Coding and agentic scores
MiniMax reports 59.0% on SWE-bench Pro, 66.0% on Terminal-Bench 2.1, 34.8% on SWE-fficiency and 74.2% on MCP Atlas as its frontier-level results.
Trained on real collaboration
Rather than assume single-turn tasks, the lab built a user simulator that reproduces requirement changes, mid-task correction and task switching, and trained and evaluated the model against it.
Long autonomous runs
Reproducing a research paper in about twelve hours, and lifting an FP8 kernel from 7.6% to 71.3% of hardware peak over roughly a day, with M3's best result arriving on its 145th submission.
What it is not for
MiniMax reports it behind Claude Opus 4.7 and GPT-5.5 on autonomous post-training work, and describes its financial-domain agent performance as only starting to be usable.

Indian languages

MiniMax M3 answers in 8 Indian languages. Choose a language from the menu beside the message box to receive replies in it.

Frequently Asked Questions

Frequently asked questions about MiniMax M3.

When was MiniMax M3 released?

MiniMax released MiniMax M3 on May 31, 2026.

Who built MiniMax M3?

MiniMax M3 is developed by MiniMax. 99Models AI connects directly to it at the provider's published rate.

How intelligent is MiniMax M3?

It is ranked 33 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does MiniMax M3 cost?

Usage costs ₹57.60 per million input tokens and ₹230.40 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is MiniMax M3 pricing in US dollars?

The provider charges $0.60 per million input tokens and $2.40 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can MiniMax M3 hold?

Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does MiniMax M3 rank for value?

It ranks 19 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does MiniMax M3 support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

Which Indian languages does MiniMax M3 support?

It answers in 8 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.

Catalog updated Sep 9, 2026