99MODELS

मिनीमॅक्स M3

नवीन

Multimodal 1M-context foundation model for agents and coding.

लाँच दिनांक: 31 मे, 2026

230.40

प्रति 10 लाख आउटपुट Tokens

इनपुट: ₹57.60 प्रति 10 लाख Tokens

प्रोव्हायडरच्या मूळ दरानुसार 0% मार्कअपसह प्रति US डॉलर ₹96 दराने रूपांतरित।

वैशिष्ट्ये

Context विंडो
10,48,576 Tokens
कमाल आउटपुट
4,71,859 Tokens
स्वीकारतो
टेक्स्ट, इमेजेस, व्हिडिओ
Reasoning
बाय डीफॉल्ट चालू
टूल वापर
होय
स्ट्रक्चर्ड आउटपुट
होय
Code एक्झिक्यूशन
नाही
इंटेलिजन्स रँक
54 पैकी #33
व्हॅल्यू रँक
54 पैकी #19

दरपत्रक

दरपत्रक
प्रति 10 lakh TokensINRUSD
इनपुट57.60$0.60
आउटपुट230.40$2.40
कॅश केलेला इनपुट5.76$0.06

बेंचमार्क

रेटिंग म्हणून नमूद केलेले नसल्यास स्कोअर टक्केवारीत आहेत. सर्व बेंचमार्क स्वतंत्रपणे तपासले जातात.

  • 92.9%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 39.0%

    HLE

    Humanity's Last Exam

  • 47.1%

    SciCode

    SciCode - scientific code generation

  • 82.9%

    IFBench

    IFBench - precise instruction following

  • 83.0%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 42.4%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 65.2%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 14.7%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 22.3%

    OSWorld 2

    OSWorld 2 - agentic computer use (partial credit)

  • 1488

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

मिनीमॅक्स M3 विषयी

MiniMax's native multimodal model with a million-token context, built with roughly 428B total and 23B activated parameters. It uses MiniMax Sparse Attention for efficient long context and is trained on mixed modalities from inception, so text, images and video are integrated semantically rather than having vision bolted on afterwards. MiniMax presents it as the first open-weight model with frontier coding, agentic reasoning and native multimodality at once.

MiniMax makes a combination claim rather than a peak-score one. Frontier coding, a million tokens of context and native multimodality are each unremarkable in a closed frontier model, and M3 is the lab's argument that one set of open weights can carry all three at the same time. What makes that affordable is MSA, the sparse attention design the team proposes: it partitions the key-value cache into blocks more precisely than the alternatives it names, and its operator uses the blocks as the outer loop so each is read exactly once with contiguous memory access.

The efficiency figures are what turn the context from a number into something usable. MiniMax reports the operator running more than four times faster than two open sparse-attention implementations, per-token compute at a million tokens of one twentieth of the previous generation, and end-to-end speedups of more than nine times in prefill and more than fifteen times in decode. Across several ablations it says MSA matched full attention on the vast majority of capabilities, which is the claim that matters most and the easiest one to get wrong.

For coding the lab reports 59.0% on SWE-bench Pro, 66.0% on Terminal-Bench 2.1, 34.8% on SWE-fficiency, 28.8% on the hard tier of a GPU kernel benchmark and 74.2% on MCP Atlas. More interesting than the numbers is the training argument underneath them. MiniMax says most code-agent training and evaluation assumes a single-turn task, which is not how anyone actually works, so it built an interactive user simulator that reproduces requirement elaboration, mid-task correction, task switching and multi-round iteration, and trained and evaluated against that instead.

Its long-run examples follow the same shape. Handed an award-winning machine learning paper and asked to reproduce it independently, M3 ran for close to twelve hours, produced 18 commits and 23 experimental figures and completed the core experiments, including reproducing the effect the paper is known for and verifying the mitigation it proposes. On an FP8 matrix-multiplication kernel with no reference implementation available to imitate, it ran for roughly 24 hours across 147 benchmark submissions and 1,959 tool calls, lifting hardware peak utilisation from 7.6% to 71.3%. MiniMax notes that most other models it tried stopped making progress within the first 30 submissions, while M3's best result arrived on submission 145, after several plateaus it worked through rather than gave up on.

The lab is explicit about where it is not yet ahead. On the post-training benchmark, where the model must autonomously synthesise data, train, evaluate and iterate on four base models inside twelve hours, it scored 0.37, behind Claude Opus 4.7 at 0.42 and GPT-5.5 at 0.39 though clearly ahead of everything else it tested. It describes its agentic performance in the financial domain as only beginning to be usable rather than solved. Weights and a technical report were promised within ten days of the post, so this is a model you could host yourself; the copy served here is MiniMax's.

लाँचवेळी MiniMax ने काय सांगितले

Three capabilities at once
MiniMax presents M3 as the first open-weight model to combine frontier coding, a million-token context and native multimodality, rather than leading on any one of them alone.
Sparse attention that scales
The lab measures its MSA operator more than four times faster than two open sparse-attention implementations, with prefill over nine times and decode over fifteen times faster at long context.
Coding and agentic scores
MiniMax reports 59.0% on SWE-bench Pro, 66.0% on Terminal-Bench 2.1, 34.8% on SWE-fficiency and 74.2% on MCP Atlas as its frontier-level results.
Trained on real collaboration
Rather than assume single-turn tasks, the lab built a user simulator that reproduces requirement changes, mid-task correction and task switching, and trained and evaluated the model against it.
Long autonomous runs
Reproducing a research paper in about twelve hours, and lifting an FP8 kernel from 7.6% to 71.3% of hardware peak over roughly a day, with M3's best result arriving on its 145th submission.
What it is not for
MiniMax reports it behind Claude Opus 4.7 and GPT-5.5 on autonomous post-training work, and describes its financial-domain agent performance as only starting to be usable.

भारतीय भाषा

मिनीमॅक्स M3 हे 8 भारतीय भाषांमध्ये उत्तरे देते। मेसेज बॉक्ससमोरील मेनूमधून भाषा निवडा आणि त्याच भाषेत उत्तर मिळवा।

Frequently Asked Questions

मिनीमॅक्स M3 बद्दल वारंवार विचारले जाणारे प्रश्न।

मिनीमॅक्स M3 कधी लाँच झाले?

MiniMax ने मिनीमॅक्स M3 मॉडेल 31 मे, 2026 रोजी लाँच केले.

मिनीमॅक्स M3 ची निर्मिती कोणी केली?

मिनीमॅक्स M3 ची निर्मिती MiniMax ने केली आहे। 99Models प्रोव्हायडरच्या मूळ दरात थेट तिच्याशी जोडते।

मिनीमॅक्स M3 किती कार्यक्षम आहे?

स्वतंत्र बेंचमार्क गुणांवर आधारित आमच्या बुद्धिमत्ता रँकिंगमध्ये 54 पैकी या Model चा क्रमांक 33 आहे। तिचे सर्व गुण वरील Benchmarks तक्त्यामध्ये पाहू शकता।

मिनीमॅक्स M3 चे दर किती आहेत?

0% मार्कअपसह दर प्रति 10 लाख इनपुट Tokens साठी ₹57.60 आणि प्रति 10 लाख आउटपुट Tokens साठी ₹230.40 आहे। कोणतेही सबस्क्रिप्शन नाही; तुम्ही वापरानुसार पेमेंट करता।

मिनीमॅक्स M3 चे अमेरिकन डॉलरमधील दर काय आहेत?

प्रोव्हायडर प्रति 10 लाख इनपुट Tokens साठी $0.60 आणि प्रति 10 लाख आउटपुट Tokens साठी $2.40 आकारतो। रुपयांचे दर प्रति अमेरिकन डॉलर ₹96 या दराने रूपांतरित केले आहेत।

मिनीमॅक्स M3 किती मोठे संभाषण लक्षात ठेवू शकते?

याची Context विंडो 10.5 lakh Tokens आहे। एकाच विनंतीमध्ये हे Model संभाषण आणि जोडलेल्या फाइल्स मिळून एवढा एकूण मजकूर वाचू शकते।

मूल्याच्या (Value) बाबतीत मिनीमॅक्स M3 चा क्रमांक कितवा आहे?

मूल्य रँकिंगमध्ये 54 मॉडेलपैकी हिचा क्रमांक 19 आहे। हे रँकिंग मॉडेलची बुद्धिमत्ता आणि Tokens च्या किमतीची तुलना करून ठरवले जाते।

मिनीमॅक्स M3 उत्तर देण्यापूर्वी विचार (Reasoning) करते का?

बाय डीफॉल्ट चालू। Reasoning उपलब्ध असल्यास, तुम्ही मेसेज कंपोजरमध्ये विचार करण्याची पातळी (effort level) निवडू शकता।

मिनीमॅक्स M3 कोणत्या भारतीय भाषांमध्ये उत्तरे देते?

हे 8 भारतीय भाषांमध्ये उत्तरे देते। मेसेज बॉक्ससमोरील मेनूमधून भाषा निवडा आणि त्याच भाषेत उत्तर मिळवा।

कॅटलॉग अपडेट: 9 सप्टें, 2026