میمو v2.5
Cheap 1M-context multimodal model; one of the highest-volume models anywhere.
تاریخ اجراء: 22 اپریل، 2026
₹192.00
فی 10 لاکھ آؤٹ پٹ Tokens
ان پٹ: ₹38.40 فی 10 لاکھ Tokens
فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔
تکنیکی تفصیلات
- Context ونڈو
- 10,50,000 Tokens
- زیادہ سے زیادہ آؤٹ پٹ
- 1,31,072 Tokens
- سپورٹ
- ٹیکسٹ, تصاویر, ویڈیو
- Reasoning
- پہلے سے آن
- ٹول کا استعمال
- ہاں
- سٹرکچرڈ آؤٹ پٹ
- ہاں
- کوڈ ایگزیکیوشن
- نہیں
- ذہانت کا درجہ
- 54 میں سے #46
- ویلیو کا درجہ
- 54 میں سے #44
قیمتیں
Benchmarks
تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔
76.3%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
25.2%
HLE
Humanity's Last Exam
43.1%
SciCode
SciCode - scientific code generation
41.7%
Terminal-Bench Hard
Terminal-Bench Hard - agentic terminal tasks
1438
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
میمو v2.5 کے بارے میں
Xiaomi's natively omnimodal model, handling text, image, video and audio in one unified architecture. It is a sparse Mixture-of-Experts with 310B total and 15B active parameters, pairing a vision encoder and an audio transformer with the language backbone, and it inherits a hybrid sliding-window attention design that cuts KV-cache storage nearly sixfold. It supports up to a million tokens of context at a very low price.
Xiaomi built this one to perceive and to act in the same pass. The language backbone is inherited from the team's earlier hybrid sliding-window model, and the vision and audio encoders - both pretrained in-house - hang off it through lightweight projectors rather than being bolted on as a separate pipeline. The lab's summary of the goal is a single model that sees, hears and acts on what it perceives, and the training schedule is built around that: text pre-training for the backbone, a projector warmup to align audio and vision with the language model, large-scale multimodal pre-training, supervised fine-tuning and agentic post-training during which the window is stretched from 32K to 256K to a million tokens, and finally reinforcement learning with multi-teacher on-policy distillation.
The agent claims are where Xiaomi puts its emphasis, and it argues them on efficiency as much as accuracy. On its internal coding evaluation it reports the model matching its own far larger Pro sibling at half the cost. On its daily-agent benchmark it reports 62.3 on the general subset and places the result at the frontier of score against token spend, which is the whole pitch: frontier-level agent behaviour without frontier-level token bills. On the public coding sets it reports 65.8 on Terminal-Bench 2.0 and 56.1 on SWE-Bench Pro.
The perception numbers are the other half. Xiaomi reports 81.0 on the chart-reasoning benchmark CharXiv, 77.9 on MMMU-Pro, 88.5 on high-resolution image understanding, 87.2 on document understanding, 87.7 on general video question answering and 64.0 on the harder video-reasoning set VideoHolmes. Its framing of those is comparative and specific: level with a leading closed model on video, level with another on multimodal agentic work, and competitive rather than leading on image and document understanding.
Where it stops shows up in the same table. On the multimodal agent split - real, messy interactions rather than single questions - Xiaomi reports 23.8, behind two of the closed models it compared against, and the number is low in absolute terms for every model on that row. The lab also positions the larger Pro model above this one for the hardest long-horizon software engineering. Read this as the omnimodal workhorse: cheap, very long-context, strong at everyday agent work and at reading what it is shown, with the hardest autonomous runs left to its bigger sibling.
لانچ کے وقت Xiaomi کا بیان
- One model across four modalities
- Vision and audio encoders pretrained in-house attach to the language backbone through lightweight projectors, so text, image, video and audio are reasoned about in a single architecture.
- Context grown during post-training
- Xiaomi extended the window from 32K to 256K to a million tokens across supervised fine-tuning and agentic post-training, rather than bolting long context on afterwards.
- Agent quality per token
- The lab reports 62.3 on the general subset of its daily-agent benchmark and places the model at the frontier of score against token spend, matching its Pro sibling's internal coding score at half the cost.
- Perception benchmarks
- Xiaomi reports 81.0 on CharXiv chart reasoning, 77.9 on MMMU-Pro, 88.5 on high-resolution images, 87.2 on documents and 87.7 on general video question answering.
- Coding agent results
- On the public sets the lab reports 65.8 on Terminal-Bench 2.0 and 56.1 on SWE-Bench Pro, alongside its internal coding evaluation.
- What it is not for
- On Xiaomi's own multimodal agent split it scores 23.8, behind two of the closed models it compared against, and the lab points at the larger Pro model for the hardest long-horizon engineering.
ہندوستانی زبانیں
میمو v2.5 5 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے زبان کا انتخاب کریں۔
Frequently Asked Questions
میمو v2.5 کے بارے میں اکثر پوچھے جانے والے سوالات۔
میمو v2.5 کب جاری ہوا تھا؟
Xiaomi نے میمو v2.5 کو 22 اپریل، 2026 کو جاری کیا۔
میمو v2.5 کس نے بنایا ہے؟
میمو v2.5 کو Xiaomi نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔
میمو v2.5 کتنا ذہین ہے؟
یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 46 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔
میمو v2.5 کا کتنا خرچ آتا ہے؟
استعمال کی لاگت ₹38.40 فی 10 لاکھ ان پٹ Tokens اور ₹192.00 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔
امریکی ڈالر میں میمو v2.5 کی قیمت کیا ہے؟
فراہم کنندہ $0.40 فی 10 لاکھ ان پٹ Tokens اور $2.00 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔
میمو v2.5 کتنی لمبی گفتگو یاد رکھ سکتا ہے؟
اس کی Context حد 10.5 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔
کیا میمو v2.5 مناسب قیمت میں بہترین کارکردگی دیتا ہے؟
یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 44 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔
کیا میمو v2.5 جواب دینے سے پہلے سوچتا ہے؟
پہلے سے آن۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔
میمو v2.5 کن ہندوستانی زبانوں میں جواب دیتا ہے؟
یہ 5 ہندوستانی زبانوں میں جواب دیتا ہے۔ میسج باکس کے ساتھ والے مینو سے اپنی زبان منتخب کریں۔
Xiaomi کے مزید ماڈلز
کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026