Gemini 2.5 Flash
Long-serving workhorse Flash model; broad multimodal input.
Released Jun 17, 2025
₹240.00
per 10 lakh output tokens
Input: ₹28.80 per 10 lakh tokens
Billed at provider rates converted at ₹96 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
83.2%
MMLU-Pro
MMLU-Pro - multitask language understanding
79.0%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
12.1%
HLE
Humanity's Last Exam
69.5%
LiveCodeBench
LiveCodeBench - contamination-free coding
98.1%
MATH-500
MATH-500 - competition mathematics, 500 problems
82.3%
AIME 2024
AIME 2024 - competition mathematics
73.3%
AIME 2025
AIME 2025 - competition mathematics
50.3%
IFBench
IFBench - precise instruction following
65.3%
Long Context
Long Context Reasoning - reasoning over long inputs
13.6%
Terminal-Bench Hard
Terminal-Bench Hard - agentic terminal tasks
About Gemini 2.5 Flash
Google's price-performance workhorse of its generation, with well-rounded capabilities across reasoning, coding, mathematics and scientific tasks. It is best suited to large-scale processing, low-latency high-volume work and agentic use cases that still need thinking. Built-in thinking can be given a token budget or switched off entirely for speed.
This is the older, cheaper, thoroughly understood option in the catalog, and that is the reason to pick it. Gemini 2.5 Flash reached general availability alongside 2.5 Pro in the post cited here, unchanged from the preview many teams had already tuned against, and it has been in production ever since. Nothing about it is going to move under a workload that already works.
The interesting part of that announcement was the pricing, not the model. Google collapsed the separate thinking and non-thinking rates into one after developers found the split confusing, raised the input price and cut the output price, and kept a single tier regardless of how large the prompt is - so a long-context request costs the same per token as a short one. That last point is easy to miss and is worth checking against any newer model before switching.
Google's own placement of it is plain: Flash-Lite for high-volume cost-efficient work, Flash for fast performance on everyday tasks, Pro for coding and highly complex tasks. Every model in the 2.5 family is a thinking model whose budget the caller sets, and on Flash that budget can be turned down to nothing, which is what makes the same endpoint usable for both a reasoning step and a bulk classification pass.
What it is not is current. Google has shipped several Flash generations since, and measures each new one as cheaper per completed task and stronger on coding, document and agentic work than the one before it. 2.5 Flash is the right answer for a pipeline whose prompts are already calibrated to it, or for a job that needs a well-known model at a low, stable rate. It is the wrong answer for anything that needs the current state of the Flash line.
What Google announced at launch
- Stable, and deliberately so
- It went generally available with no change from the preview many teams had already built against, which is the whole argument for a model at this age: nothing shifts under a working prompt.
- One rate, one tier
- Google merged the separate thinking and non-thinking prices into a single rate after developers found the split confusing, and kept one price tier regardless of how large the input is.
- A thinking budget you set
- Every 2.5 model reasons before answering with a budget the caller controls, and on Flash that budget can be set to zero, which makes one endpoint serve both reasoning and bulk work.
- Where Google put it
- The lab's own ladder reads Flash-Lite for high-volume cost-efficient tasks, Flash for fast performance on everyday tasks, and Pro for coding and highly complex tasks.
- What it is not for
- It is a generation behind. Google measures its newer Flash models as cheaper per completed task and stronger on coding, document and agentic work, and points new builds at those instead.
Indian languages
Gemini 2.5 Flash answers in 14 Indian languages. Choose a language from the menu beside the message box to receive replies in it.
Frequently Asked Questions
Frequently asked questions about Gemini 2.5 Flash.
When was Gemini 2.5 Flash released?
Google released Gemini 2.5 Flash on Jun 17, 2025.
Who built Gemini 2.5 Flash?
Gemini 2.5 Flash is developed by Google. 99Models AI connects directly to it at the provider's published rate.
How intelligent is Gemini 2.5 Flash?
It is ranked 50 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does Gemini 2.5 Flash cost?
Usage costs ₹28.80 per million input tokens and ₹240.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is Gemini 2.5 Flash pricing in US dollars?
The provider charges $0.30 per million input tokens and $2.50 per million output tokens. Rupee rates are converted at ₹96 per US dollar.
How long a conversation can Gemini 2.5 Flash hold?
Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does Gemini 2.5 Flash rank for value?
It ranks 52 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does Gemini 2.5 Flash support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
Which Indian languages does Gemini 2.5 Flash support?
It answers in 14 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.
More models from Google
Catalog updated Sep 9, 2026