99MODELS

मिमो v2.5

Cheap 1M-context multimodal model; one of the highest-volume models anywhere.

रिलिज मिति: 2026 अप्रिल 22

192.00

प्रति 10 लाख आउटपुट Tokens

इनपुट: ₹38.40 प्रति 10 लाख Tokens

प्रति अमेरिकी डलर ₹96 मा 0% मार्कअपका साथ प्रदायककै दरमा गणना गरिन्छ।

विवरण

Context विन्डो
10,50,000 Tokens
अधिकतम आउटपुट
1,31,072 Tokens
स्वीकार गर्छ
टेक्स्ट, तस्बिरहरू, भिडियो
Reasoning
सुरुमै चालु
टुल प्रयोग
संरचित आउटपुट
कोड कार्यान्वयन
छैन
इन्टेलिजेन्स र्‍याङ्क
54 मध्ये #46
भ्याल्यू र्‍याङ्क
54 मध्ये #44

मूल्य

मूल्य
प्रति 10 lakh TokensINRUSD
इनपुट38.40$0.40
आउटपुट192.00$2.00
क्यास गरिएको इनपुट7.68$0.08

बेन्चमार्क

रेटिङ बाहेकका सबै स्कोरहरू प्रतिशतमा छन्। सबै बेन्चमार्कहरू स्वतन्त्र रूपमा मापन गरिएका हुन्।

  • 76.3%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 25.2%

    HLE

    Humanity's Last Exam

  • 43.1%

    SciCode

    SciCode - scientific code generation

  • 41.7%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 1438

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

मिमो v2.5 को बारेमा

Xiaomi's natively omnimodal model, handling text, image, video and audio in one unified architecture. It is a sparse Mixture-of-Experts with 310B total and 15B active parameters, pairing a vision encoder and an audio transformer with the language backbone, and it inherits a hybrid sliding-window attention design that cuts KV-cache storage nearly sixfold. It supports up to a million tokens of context at a very low price.

Xiaomi built this one to perceive and to act in the same pass. The language backbone is inherited from the team's earlier hybrid sliding-window model, and the vision and audio encoders - both pretrained in-house - hang off it through lightweight projectors rather than being bolted on as a separate pipeline. The lab's summary of the goal is a single model that sees, hears and acts on what it perceives, and the training schedule is built around that: text pre-training for the backbone, a projector warmup to align audio and vision with the language model, large-scale multimodal pre-training, supervised fine-tuning and agentic post-training during which the window is stretched from 32K to 256K to a million tokens, and finally reinforcement learning with multi-teacher on-policy distillation.

The agent claims are where Xiaomi puts its emphasis, and it argues them on efficiency as much as accuracy. On its internal coding evaluation it reports the model matching its own far larger Pro sibling at half the cost. On its daily-agent benchmark it reports 62.3 on the general subset and places the result at the frontier of score against token spend, which is the whole pitch: frontier-level agent behaviour without frontier-level token bills. On the public coding sets it reports 65.8 on Terminal-Bench 2.0 and 56.1 on SWE-Bench Pro.

The perception numbers are the other half. Xiaomi reports 81.0 on the chart-reasoning benchmark CharXiv, 77.9 on MMMU-Pro, 88.5 on high-resolution image understanding, 87.2 on document understanding, 87.7 on general video question answering and 64.0 on the harder video-reasoning set VideoHolmes. Its framing of those is comparative and specific: level with a leading closed model on video, level with another on multimodal agentic work, and competitive rather than leading on image and document understanding.

Where it stops shows up in the same table. On the multimodal agent split - real, messy interactions rather than single questions - Xiaomi reports 23.8, behind two of the closed models it compared against, and the number is low in absolute terms for every model on that row. The lab also positions the larger Pro model above this one for the hardest long-horizon software engineering. Read this as the omnimodal workhorse: cheap, very long-context, strong at everyday agent work and at reading what it is shown, with the hardest autonomous runs left to its bigger sibling.

लन्चको समयमा Xiaomi ले के भन्यो

One model across four modalities
Vision and audio encoders pretrained in-house attach to the language backbone through lightweight projectors, so text, image, video and audio are reasoned about in a single architecture.
Context grown during post-training
Xiaomi extended the window from 32K to 256K to a million tokens across supervised fine-tuning and agentic post-training, rather than bolting long context on afterwards.
Agent quality per token
The lab reports 62.3 on the general subset of its daily-agent benchmark and places the model at the frontier of score against token spend, matching its Pro sibling's internal coding score at half the cost.
Perception benchmarks
Xiaomi reports 81.0 on CharXiv chart reasoning, 77.9 on MMMU-Pro, 88.5 on high-resolution images, 87.2 on documents and 87.7 on general video question answering.
Coding agent results
On the public sets the lab reports 65.8 on Terminal-Bench 2.0 and 56.1 on SWE-Bench Pro, alongside its internal coding evaluation.
What it is not for
On Xiaomi's own multimodal agent split it scores 23.8, behind two of the closed models it compared against, and the lab points at the larger Pro model for the hardest long-horizon engineering.

भारतीय भाषाहरू

मिमो v2.5 ले 5 भारतीय भाषाहरूमा जवाफ दिन्छ। सन्देश बाकस छेउको मेनुबाट भाषा छान्नुहोस् र सोही भाषामा जवाफ पाउनुहोस्।

Frequently Asked Questions

मिमो v2.5 सम्बन्धी प्रायः सोधिने प्रश्नहरू।

मिमो v2.5 कहिले रिलिज भएको हो?

Xiaomi ले मिमो v2.5 लाई 2026 अप्रिल 22 मा सार्वजनिक गरेको हो।

मिमो v2.5 कसले बनाएको हो?

मिमो v2.5 लाई Xiaomi ले बनाएको हो। 99Models AI ले प्रदायककै दरमा सिधै जोड्दछ।

मिमो v2.5 कत्तिको सक्षम र बुद्धिमानी छ?

यो हाम्रो बौद्धिकता श्रेणीकरणमा 54 च्याट Models मध्ये 46 स्थानमा छ। यसको पूर्ण अङ्क माथिको Benchmarks तालिकामा हेर्न सकिन्छ।

मिमो v2.5 को लागत कति पर्छ?

यसमा 0% मार्कअपका साथ प्रति 10 लाख इनपुट Tokens को ₹38.40 र आउटपुटको ₹192.00 लाग्छ। कुनै सदस्यता छैन; तपाईंले प्रयोग गरेअनुसार मात्र भुक्तानी गर्नुहुन्छ।

डलरमा मिमो v2.5 को API मूल्य कति हो?

प्रदायकले प्रति 10 लाख इनपुट Tokens को $0.40 र आउटपुट Tokens को $2.00 शुल्क लिन्छ। यस पृष्ठका दरहरू प्रति अमेरिकी डलर ₹96 मा रूपान्तरण गरिएका हुन्।

मिमो v2.5 ले कति लामो कुराकानी सम्झन सक्छ?

यसको Context विन्डो 10.5 lakh Tokens हो। यसले एकल अनुरोधमा प्रक्रिया गर्न सक्ने कुराकानी र संलग्न फाइलहरूको कुल क्षमता यही हो।

के मिमो v2.5 लागत अनुसार उत्कृष्ट छ?

मूल्य र गुणस्तरको आधारमा यो 54 Models मध्ये 44 स्थानमा छ। यसले बौद्धिकता र Token लागतको तुलना गर्दछ।

के मिमो v2.5 ले जवाफ दिनुअघि विचार गर्छ?

सुरुमै चालु। Reasoning उपलब्ध भएको ठाउँमा तपाईंले सिधै सन्देश बक्समा सोच्ने क्षमता समायोजन गर्न सक्नुहुन्छ।

मिमो v2.5 ले कुन-कुन भारतीय भाषाहरूमा जवाफ दिन्छ?

यसले 5 भारतीय भाषाहरूमा जवाफ दिन्छ। सन्देश बाकस छेउको मेनुबाट भाषा छान्नुहोस् र सोही भाषामा जवाफ पाउनुहोस्।

Xiaomi का अन्य Models

क्याटलग अद्यावधिक गरिएको मिति: 2026 सेप्टेम्बर 9