Gemini 3.7 Flash
NewFast multimodal model for agentic workflows, coding and multi-step reasoning.
Released Aug 13, 2026
₹360.00
per 10 lakh output tokens
Input: ₹72.00 per 10 lakh tokens
Billed at provider rates converted at ₹96 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
94.5%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
47.9%
HLE
Humanity's Last Exam
57.2%
SciCode
SciCode - scientific code generation
81.7%
Long Context
Long Context Reasoning - reasoning over long inputs
85.8%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
43.6%
FrontierCode
FrontierCode - long-horizon production coding tasks
84.6%
ARC-AGI-2
ARC-AGI-2 - abstract reasoning on novel puzzles
1587
WebDev Arena
WebDev Arena - head-to-head web-app builds, Elo rating
About Gemini 3.7 Flash
The next iteration in Google's Gemini 3 series of natively multimodal reasoning models, and what Google calls its most intelligent workhorse model yet for coding and agents. Algorithmic improvements to its core reasoning give substantial gains across software engineering, knowledge work and web development, with better debugging and higher first-pass accuracy on production-ready code than 3.6 Flash. It reads text, images, audio and video, and always reasons -- the thinking level is adjustable but cannot be switched off.
Flash is the middle rung of the Gemini ladder: more capable than the Flash-Lite models Google sells for throughput, cheaper and faster than the Pro line it sells for the hardest reasoning. 3.7 Flash is the top of that middle rung, and Google shipped it three weeks after 3.6 Flash rather than waiting for a generation boundary, describing it as the result of developer feedback and of algorithmic work it expects to carry into later models.
The gains Google publishes are concentrated in the places a workhorse model actually earns its keep. On FrontierCode 1.1 Main, which scores production code quality, the lab reports 43.6% against 34.4% for 3.6 Flash, and on the long-horizon software engineering set DeepSWE v1.1, 65.3% against 49.0%. On the head-to-head web development evaluation Google cites, it moves the Elo from 1538 to 1588 and generates feature-complete layouts in fewer prompts, including from a screenshot or a full design system as the reference. On agentic terminal work Google reports 85.8% against 78.0%.
Two of the largest jumps are in knowledge work rather than code. GDP.pdf, an evaluation of expert document comprehension, goes from 22.0% to 34.0%, and AutomationBench, which asks a model to complete real business workflows end to end, nearly doubles from 17.0% to 30.4%. Google also reports 97.0% on the eight-needle long context set at 128k against 91.8% for 3.6 Flash, which is the number to weigh against a million-token window.
Google spends as much of the post on behaviour as on scores. It says the model adapts to roadblocks, asks for clarification when a request is ambiguous, follows instructions more closely and puts more effort into multi-step planning and tool calls - which it frames as less manual oversight and fewer retries rather than as a benchmark result. The price is worth reading carefully too: Google launched it at half of what 3.6 Flash originally cost per million tokens, but that rate is introductory and runs to 31 December 2026, after which the lab has said input and output both double.
Where it stops is visible in Google's own table. On Terminal-bench 3.0, its broadest general-agent evaluation, 3.7 Flash scores 14.9%, and on OSWorld-2.0 computer use 47.9% - both far below the rows this model wins, and a reminder that autonomous desktop work is still early for every model in that table. The release also ships with tightened safeguards in the chemical, biological, radiological, nuclear and cyber-offence domains.
What Google announced at launch
- Coding and issue resolution
- Google reports 43.6% against 34.4% for 3.6 Flash on FrontierCode 1.1 Main, and 65.3% against 49.0% on the long-horizon DeepSWE v1.1 set, with better debugging and higher first-pass accuracy.
- Web development from a reference
- The lab reports an Elo of 1588 against 1538 for 3.6 Flash on the web development head-to-head it cites, with high design adherence when given a screenshot, an image or a full design system as the input.
- Documents and business workflows
- On the GDP.pdf document comprehension evaluation Google reports 34.0% against 22.0%, and on AutomationBench, which runs real business workflows end to end, 30.4% against 17.0%.
- Fewer retries in practice
- Google describes more disciplined execution: adapting to roadblocks, clarifying ambiguous intent, closer instruction following and more effort spent on multi-step planning and tool calls.
- Introductory pricing
- The launch rate is an introductory one that Google says expires on 31 December 2026, after which both the input and the output price double.
- What it is not for
- Google's own table puts it at 14.9% on Terminal-bench 3.0 and 47.9% on OSWorld-2.0, so open-ended computer use is not where this model is strong. It also ships with tightened CBRN and cyber-offence safeguards.
Indian languages
Gemini 3.7 Flash answers in 15 Indian languages. Choose a language from the menu beside the message box to receive replies in it.
Frequently Asked Questions
Frequently asked questions about Gemini 3.7 Flash.
When was Gemini 3.7 Flash released?
Google released Gemini 3.7 Flash on Aug 13, 2026.
Who built Gemini 3.7 Flash?
Gemini 3.7 Flash is developed by Google. 99Models AI connects directly to it at the provider's published rate.
How intelligent is Gemini 3.7 Flash?
It is ranked 20 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does Gemini 3.7 Flash cost?
Usage costs ₹72.00 per million input tokens and ₹360.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is Gemini 3.7 Flash pricing in US dollars?
The provider charges $0.75 per million input tokens and $3.75 per million output tokens. Rupee rates are converted at ₹96 per US dollar.
How long a conversation can Gemini 3.7 Flash hold?
Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does Gemini 3.7 Flash rank for value?
It ranks 8 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does Gemini 3.7 Flash support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
Which Indian languages does Gemini 3.7 Flash support?
It answers in 15 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.
More models from Google
Catalog updated Sep 9, 2026