99MODELS

Gemini 3.7 Flash

New

Fast multimodal model for agentic workflows, coding and multi-step reasoning.

Released Aug 13, 2026

360.00

per 10 lakh output tokens

Input: ₹72.00 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
10,48,576 tokens
Max output
65,536 tokens
Accepts
Text, Images, PDF files, Audio, Video
Reasoning
On by default
Effort levels
low, medium, high
Tool use
Yes
Structured output
Yes
Code execution
Yes
Intelligence rank
#20 of 54
Value rank
#8 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input72.00$0.75
Output360.00$3.75
Cached input7.20$0.08

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 94.5%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 47.9%

    HLE

    Humanity's Last Exam

  • 57.2%

    SciCode

    SciCode - scientific code generation

  • 81.7%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 85.8%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 43.6%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 84.6%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 1587

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

About Gemini 3.7 Flash

The next iteration in Google's Gemini 3 series of natively multimodal reasoning models, and what Google calls its most intelligent workhorse model yet for coding and agents. Algorithmic improvements to its core reasoning give substantial gains across software engineering, knowledge work and web development, with better debugging and higher first-pass accuracy on production-ready code than 3.6 Flash. It reads text, images, audio and video, and always reasons -- the thinking level is adjustable but cannot be switched off.

Flash is the middle rung of the Gemini ladder: more capable than the Flash-Lite models Google sells for throughput, cheaper and faster than the Pro line it sells for the hardest reasoning. 3.7 Flash is the top of that middle rung, and Google shipped it three weeks after 3.6 Flash rather than waiting for a generation boundary, describing it as the result of developer feedback and of algorithmic work it expects to carry into later models.

The gains Google publishes are concentrated in the places a workhorse model actually earns its keep. On FrontierCode 1.1 Main, which scores production code quality, the lab reports 43.6% against 34.4% for 3.6 Flash, and on the long-horizon software engineering set DeepSWE v1.1, 65.3% against 49.0%. On the head-to-head web development evaluation Google cites, it moves the Elo from 1538 to 1588 and generates feature-complete layouts in fewer prompts, including from a screenshot or a full design system as the reference. On agentic terminal work Google reports 85.8% against 78.0%.

Two of the largest jumps are in knowledge work rather than code. GDP.pdf, an evaluation of expert document comprehension, goes from 22.0% to 34.0%, and AutomationBench, which asks a model to complete real business workflows end to end, nearly doubles from 17.0% to 30.4%. Google also reports 97.0% on the eight-needle long context set at 128k against 91.8% for 3.6 Flash, which is the number to weigh against a million-token window.

Google spends as much of the post on behaviour as on scores. It says the model adapts to roadblocks, asks for clarification when a request is ambiguous, follows instructions more closely and puts more effort into multi-step planning and tool calls - which it frames as less manual oversight and fewer retries rather than as a benchmark result. The price is worth reading carefully too: Google launched it at half of what 3.6 Flash originally cost per million tokens, but that rate is introductory and runs to 31 December 2026, after which the lab has said input and output both double.

Where it stops is visible in Google's own table. On Terminal-bench 3.0, its broadest general-agent evaluation, 3.7 Flash scores 14.9%, and on OSWorld-2.0 computer use 47.9% - both far below the rows this model wins, and a reminder that autonomous desktop work is still early for every model in that table. The release also ships with tightened safeguards in the chemical, biological, radiological, nuclear and cyber-offence domains.

What Google announced at launch

Coding and issue resolution
Google reports 43.6% against 34.4% for 3.6 Flash on FrontierCode 1.1 Main, and 65.3% against 49.0% on the long-horizon DeepSWE v1.1 set, with better debugging and higher first-pass accuracy.
Web development from a reference
The lab reports an Elo of 1588 against 1538 for 3.6 Flash on the web development head-to-head it cites, with high design adherence when given a screenshot, an image or a full design system as the input.
Documents and business workflows
On the GDP.pdf document comprehension evaluation Google reports 34.0% against 22.0%, and on AutomationBench, which runs real business workflows end to end, 30.4% against 17.0%.
Fewer retries in practice
Google describes more disciplined execution: adapting to roadblocks, clarifying ambiguous intent, closer instruction following and more effort spent on multi-step planning and tool calls.
Introductory pricing
The launch rate is an introductory one that Google says expires on 31 December 2026, after which both the input and the output price double.
What it is not for
Google's own table puts it at 14.9% on Terminal-bench 3.0 and 47.9% on OSWorld-2.0, so open-ended computer use is not where this model is strong. It also ships with tightened CBRN and cyber-offence safeguards.

Indian languages

Gemini 3.7 Flash answers in 15 Indian languages. Choose a language from the menu beside the message box to receive replies in it.

Frequently Asked Questions

Frequently asked questions about Gemini 3.7 Flash.

When was Gemini 3.7 Flash released?

Google released Gemini 3.7 Flash on Aug 13, 2026.

Who built Gemini 3.7 Flash?

Gemini 3.7 Flash is developed by Google. 99Models AI connects directly to it at the provider's published rate.

How intelligent is Gemini 3.7 Flash?

It is ranked 20 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does Gemini 3.7 Flash cost?

Usage costs ₹72.00 per million input tokens and ₹360.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is Gemini 3.7 Flash pricing in US dollars?

The provider charges $0.75 per million input tokens and $3.75 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can Gemini 3.7 Flash hold?

Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does Gemini 3.7 Flash rank for value?

It ranks 8 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does Gemini 3.7 Flash support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

Which Indian languages does Gemini 3.7 Flash support?

It answers in 15 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.

More models from Google

Catalog updated Sep 9, 2026