99MODELS

जेमा 4 26B A4B

Sparse 26B Gemma with 4B active parameters; cheapest vision-capable Gemma.

रिलीज़: 3 अप्रैल 2026

57.60

प्रति 10 लाख आउटपुट Tokens

इनपुट: ₹14.40 प्रति 10 लाख Tokens

प्रदाता की आधिकारिक दर पर बिलिंग, ₹96 प्रति डॉलर पर कनवर्ट, 0% मार्कअप के साथ।

स्पेसिफिकेशन्स

Context विंडो
2,62,144 Tokens
अधिकतम आउटपुट
16,384 Tokens
स्वीकार्य इनपुट
टेक्स्ट, इमेज, वीडियो
Reasoning
वैकल्पिक
टूल का उपयोग
हाँ
स्ट्रक्चर्ड आउटपुट
हाँ
कोड एग्जीक्यूशन
नहीं
पैरामीटर
26B A4B MoE
इंटेलिजेंस रैंक
54 में से #48
वैल्यू रैंक
54 में से #31

दरें

दरें
प्रति 10 lakh TokensINRUSD
इनपुट14.40$0.15
आउटपुट57.60$0.60
कैश्ड इनपुट4.80$0.05

Benchmarks

स्कोर प्रतिशत में हैं जब तक कि कोई रेटिंग न दी गई हो। सभी Benchmark स्वतंत्र रूप से मापे गए हैं।

  • 79.2%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 19.3%

    HLE

    Humanity's Last Exam

  • 72.4%

    IFBench

    IFBench - precise instruction following

  • 65.7%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 13.6%

    Terminal-Bench Hard

    Terminal-Bench Hard - agentic terminal tasks

  • 39.0%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

जेमा 4 26B A4B के बारे में

The sparse member of the Gemma 4 family, focused on latency: it activates fewer than 4B of its 26B parameters per token to deliver exceptionally fast tokens per second. It keeps the multimodal text-and-image input and 256K context of its dense sibling, and Google likewise reports it outcompeting models twenty times its size. The cheapest vision-capable Gemma here.

This is the sparse member of the Gemma 4 family, and its whole design is a trade of stored parameters for spent ones. Google says it activates only 3.8 billion of its total parameters on each token, which is what produces the tokens-per-second it was built for, while the weights that sit idle still contribute the knowledge the model was trained with. The name carries the arithmetic: 26B total, roughly 4B active.

Like the rest of Gemma 4 it is published under the Apache 2.0 licence, so the weights can be downloaded, inspected, fine-tuned and shipped in a commercial product with no gate in front of them. Google sized this member to run and fine-tune on hardware people actually have - a single 80GB accelerator unquantized, consumer GPUs quantized - so the choice between running it yourself and calling it here is a real one, decided by whether the data has to stay on your machine.

The reason to take the sparse model seriously is how little it gives up. Against the dense 31B, Google reports 82.6% against 85.2% on the multilingual MMMLU, 73.8% against 76.9% on MMMU Pro, 88.3% against 89.2% on AIME 2026, 77.1% against 80.0% on LiveCodeBench v6, 82.3% against 84.3% on GPQA Diamond, and 85.5% against 86.4% on the retail split of tau2-bench. A few points across the board, for a model activating roughly an eighth as many parameters per token.

It inherits the family's build features rather than a reduced set: function calling, structured JSON output and native system instructions for agent work, native image and video input at variable resolutions with optical character recognition and chart reading, training across more than 140 languages, and a 256K context window.

What it is not is the top of its own family. The dense 31B is ahead on every row Google published, and the gap widens on the hardest ones, so the sparse model is the choice when throughput or cost per token decides the workload rather than the last few points of quality. Audio input, as with the 31B, belongs to the two edge sizes.

लॉन्च के समय Google ने क्या कहा

Sparse by design
Google says it activates only 3.8 billion of its 26 billion parameters per token, which is where the tokens-per-second comes from while the idle weights still carry what the model knows.
Open weights, Apache 2.0
The weights are published under a permissive licence and can be downloaded, fine-tuned and shipped commercially. Running it yourself is a genuine alternative to calling it here.
Close to the dense model
Google reports 82.3% against 84.3% for the 31B on GPQA Diamond, 88.3% against 89.2% on AIME 2026, and 85.5% against 86.4% on the retail split of tau2-bench.
Same build features
It carries the family set rather than a cut-down one: function calling, structured JSON, native system instructions, image and video input with OCR and chart reading, and 140-plus languages.
What it is not for
The dense 31B leads on every row Google published, and the gap grows on the hardest ones. Audio input belongs to the two edge sizes, and the context window is 256K rather than a million.

Frequently Asked Questions

जेमा 4 26B A4B के बारे में अक्सर पूछे जाने वाले सवाल।

जेमा 4 26B A4B कब रिलीज़ हुआ था?

Google ने जेमा 4 26B A4B को 3 अप्रैल 2026 को रिलीज़ किया था।

जेमा 4 26B A4B को किसने बनाया है?

जेमा 4 26B A4B को AI लैब Google ने बनाया है। 99Models इसे सीधे प्रदाता की आधिकारिक दर पर उपलब्ध कराता है।

जेमा 4 26B A4B कितना समझदार है?

इंटेलीजेंस रैंकिंग में यह 54 चैट Models में से 48 स्थान पर है, जो बेंचमार्क स्कोर पर आधारित है। इसके पूरे स्कोर ऊपर Benchmarks पैनल में देखे जा सकते हैं।

जेमा 4 26B A4B का उपयोग करने का क्या खर्च है?

10 लाख इनपुट Tokens के लिए ₹14.40 और 10 लाख आउटपुट Tokens के लिए ₹57.60, बिना किसी अतिरिक्त मार्कअप के प्रदाता की दर पर। कोई सब्सक्रिप्शन नहीं है; आप केवल अपने उपयोग का भुगतान करते हैं।

डॉलर में जेमा 4 26B A4B API की कीमत क्या है?

प्रदाता 10 लाख इनपुट Tokens के लिए $0.15 और 10 लाख आउटपुट Tokens के लिए $0.60 चार्ज करता है। इस पेज पर रुपये की दरें ₹96 प्रति डॉलर के हिसाब से बदली गई हैं।

जेमा 4 26B A4B कितनी लंबी बातचीत याद रख सकता है?

इसकी Context विंडो 2.6 lakh Tokens है। यानी यह एक रिक्वेस्ट में पिछली बातचीत और अटैच की गई फ़ाइलों को मिलाकर इतना टेक्स्ट पढ़ सकता है।

क्या जेमा 4 26B A4B पैसे के लिहाज से किफ़ायती है?

वैल्यू रैंकिंग में यह 54 Models में से 31 स्थान पर है। यह रैंकिंग परफ़ॉर्मेंस और टोकन की वास्तविक कीमत की तुलना करके तय की जाती है।

क्या जेमा 4 26B A4B जवाब देने से पहले सोच-विचार (Reasoning) करता है?

वैकल्पिक। जहां सोचने की क्षमता उपलब्ध है, वहां आप मैसेज कंपोज़र में सीधे Reasoning का स्तर तय कर सकते हैं।

Google के अन्य Models

कैटलॉग अपडेट: 9 सित॰ 2026