جیما 4 26B A4B
Sparse 26B Gemma with 4B active parameters; cheapest vision-capable Gemma.
تاریخ اجراء: 3 اپریل، 2026
₹57.60
فی 10 لاکھ آؤٹ پٹ Tokens
ان پٹ: ₹14.40 فی 10 لاکھ Tokens
فراہم کنندہ کے نرخ کے مطابق ₹96 فی امریکی ڈالر پر تبدیل شدہ، بغیر کسی اضافی مارک اپ کے۔
تکنیکی تفصیلات
- Context ونڈو
- 2,62,144 Tokens
- زیادہ سے زیادہ آؤٹ پٹ
- 16,384 Tokens
- سپورٹ
- ٹیکسٹ, تصاویر, ویڈیو
- Reasoning
- اختیاری
- ٹول کا استعمال
- ہاں
- سٹرکچرڈ آؤٹ پٹ
- ہاں
- کوڈ ایگزیکیوشن
- نہیں
- پیرامیٹرز
- 26B A4B MoE
- ذہانت کا درجہ
- 54 میں سے #48
- ویلیو کا درجہ
- 54 میں سے #31
قیمتیں
Benchmarks
تمام اسکورز فیصد میں ہیں، ماسوائے جہاں ریٹنگ درج ہو۔ یہ تمام پیمائشیں آزادانہ طور پر کی گئی ہیں۔
79.2%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
19.3%
HLE
Humanity's Last Exam
72.4%
IFBench
IFBench - precise instruction following
65.7%
Long Context
Long Context Reasoning - reasoning over long inputs
13.6%
Terminal-Bench Hard
Terminal-Bench Hard - agentic terminal tasks
39.0%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
جیما 4 26B A4B کے بارے میں
The sparse member of the Gemma 4 family, focused on latency: it activates fewer than 4B of its 26B parameters per token to deliver exceptionally fast tokens per second. It keeps the multimodal text-and-image input and 256K context of its dense sibling, and Google likewise reports it outcompeting models twenty times its size. The cheapest vision-capable Gemma here.
This is the sparse member of the Gemma 4 family, and its whole design is a trade of stored parameters for spent ones. Google says it activates only 3.8 billion of its total parameters on each token, which is what produces the tokens-per-second it was built for, while the weights that sit idle still contribute the knowledge the model was trained with. The name carries the arithmetic: 26B total, roughly 4B active.
Like the rest of Gemma 4 it is published under the Apache 2.0 licence, so the weights can be downloaded, inspected, fine-tuned and shipped in a commercial product with no gate in front of them. Google sized this member to run and fine-tune on hardware people actually have - a single 80GB accelerator unquantized, consumer GPUs quantized - so the choice between running it yourself and calling it here is a real one, decided by whether the data has to stay on your machine.
The reason to take the sparse model seriously is how little it gives up. Against the dense 31B, Google reports 82.6% against 85.2% on the multilingual MMMLU, 73.8% against 76.9% on MMMU Pro, 88.3% against 89.2% on AIME 2026, 77.1% against 80.0% on LiveCodeBench v6, 82.3% against 84.3% on GPQA Diamond, and 85.5% against 86.4% on the retail split of tau2-bench. A few points across the board, for a model activating roughly an eighth as many parameters per token.
It inherits the family's build features rather than a reduced set: function calling, structured JSON output and native system instructions for agent work, native image and video input at variable resolutions with optical character recognition and chart reading, training across more than 140 languages, and a 256K context window.
What it is not is the top of its own family. The dense 31B is ahead on every row Google published, and the gap widens on the hardest ones, so the sparse model is the choice when throughput or cost per token decides the workload rather than the last few points of quality. Audio input, as with the 31B, belongs to the two edge sizes.
لانچ کے وقت Google کا بیان
- Sparse by design
- Google says it activates only 3.8 billion of its 26 billion parameters per token, which is where the tokens-per-second comes from while the idle weights still carry what the model knows.
- Open weights, Apache 2.0
- The weights are published under a permissive licence and can be downloaded, fine-tuned and shipped commercially. Running it yourself is a genuine alternative to calling it here.
- Close to the dense model
- Google reports 82.3% against 84.3% for the 31B on GPQA Diamond, 88.3% against 89.2% on AIME 2026, and 85.5% against 86.4% on the retail split of tau2-bench.
- Same build features
- It carries the family set rather than a cut-down one: function calling, structured JSON, native system instructions, image and video input with OCR and chart reading, and 140-plus languages.
- What it is not for
- The dense 31B leads on every row Google published, and the gap grows on the hardest ones. Audio input belongs to the two edge sizes, and the context window is 256K rather than a million.
Frequently Asked Questions
جیما 4 26B A4B کے بارے میں اکثر پوچھے جانے والے سوالات۔
جیما 4 26B A4B کب جاری ہوا تھا؟
Google نے جیما 4 26B A4B کو 3 اپریل، 2026 کو جاری کیا۔
جیما 4 26B A4B کس نے بنایا ہے؟
جیما 4 26B A4B کو Google نے تیار کیا ہے۔ 99Models AI فراہم کنندہ کے اصل نرخ پر اس سے براہ راست جوڑتا ہے۔
جیما 4 26B A4B کتنا ذہین ہے؟
یہ ذہانت کی درجہ بندی میں 54 چیٹ ماڈلز میں سے 48 نمبر پر ہے۔ اس کے تمام بینچ مارک اسکور اوپر والے ٹیبل میں دیکھے جا سکتے ہیں۔
جیما 4 26B A4B کا کتنا خرچ آتا ہے؟
استعمال کی لاگت ₹14.40 فی 10 لاکھ ان پٹ Tokens اور ₹57.60 فی 10 لاکھ آؤٹ پٹ Tokens ہے، بغیر کسی اضافی فیس کے۔ کوئی سبسکرپشن نہیں؛ صرف استعمال کی ادائیگی کریں۔
امریکی ڈالر میں جیما 4 26B A4B کی قیمت کیا ہے؟
فراہم کنندہ $0.15 فی 10 لاکھ ان پٹ Tokens اور $0.60 فی 10 لاکھ آؤٹ پٹ Tokens لیتا ہے۔ روپے کے نرخ ₹96 فی امریکی ڈالر کے حساب سے تبدیل کیے گئے ہیں۔
جیما 4 26B A4B کتنی لمبی گفتگو یاد رکھ سکتا ہے؟
اس کی Context حد 2.6 lakh Tokens ہے، یعنی وہ تمام متن اور فائلیں جو یہ ایک ہی میسج میں پڑھ سکتا ہے۔
کیا جیما 4 26B A4B مناسب قیمت میں بہترین کارکردگی دیتا ہے؟
یہ بہترین قیمت کی درجہ بندی میں 54 ماڈلز میں سے 31 نمبر پر ہے، جس میں ذہانت اور ٹوکن کی لاگت کا موازنہ کیا گیا ہے۔
کیا جیما 4 26B A4B جواب دینے سے پہلے سوچتا ہے؟
اختیاری۔ جہاں Reasoning کی سہولت موجود ہو، وہاں آپ میسج باکس میں سوچنے کی سطح خود طے کر سکتے ہیں۔
Google کے مزید ماڈلز
کیٹلاگ اپ ڈیٹ: 9 ستمبر، 2026