99MODELS

Gemini 3.6 Flash

Previous Flash generation; strong general-purpose multimodal model.

Released Jul 21, 2026

396.00

per 10 lakh output tokens

Input: ₹79.20 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
10,48,576 tokens
Max output
65,536 tokens
Accepts
Text, Images, PDF files, Audio, Video
Reasoning
On by default
Effort levels
minimal, low, medium, high
Tool use
Yes
Structured output
Yes
Code execution
Yes
Intelligence rank
#28 of 54
Value rank
#17 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input79.20$0.83
Output396.00$4.13
Cached input7.92$0.08

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 92.8%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 40.8%

    HLE

    Humanity's Last Exam

  • 53.4%

    SciCode

    SciCode - scientific code generation

  • 80.0%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 77.5%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 34.4%

    FrontierCode

    FrontierCode - long-horizon production coding tasks

  • 60.4%

    ARC-AGI-2

    ARC-AGI-2 - abstract reasoning on novel puzzles

  • 1539

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

About Gemini 3.6 Flash

A high-efficiency Gemini providing sustained frontier-level intelligence for real-world tasks at higher speed and lower cost. Google built it for the agentic era, and it excels at code generation, agentic execution and spatial reasoning while using about 17 percent fewer output tokens than the previous Flash generation at better quality. Same four input modalities and million-token context as 3.7 Flash.

Google released 3.6 Flash as the workhorse of a three-model announcement: the Flash tier for general production work, a cheaper 3.5 Flash-Lite underneath it for throughput, and a restricted 3.5 Flash Cyber for vulnerability work that Google keeps to governments and trusted partners. Its stated design goal was not a higher peak score but a lower cost per completed agentic task, which is a different thing from a lower price per token.

That is why the efficiency claim leads the post. Google reports the model spending fewer output tokens than 3.5 Flash for better answers, and taking fewer reasoning steps and tool calls to finish a multi-step workflow, while charging less per token than the model it replaced. A cheaper token that is spent twice as often is not cheaper, and the lab measures the difference at the task level.

On capability the lab reports higher precision with fewer unwanted code edits and fewer execution loops on DeepSWE, at 49% against 37% for 3.5 Flash; a large jump in machine learning research work on MLE Bench, at 63.9% against 49.7%; better computer use on OSWorld-Verified, at 83.0% against 78.4%; and better knowledge work on the GDPval-AA v2 Elo, at 1421 against 1349. Computer use also became a built-in client-side tool with this release rather than something the caller has to assemble.

The multimodal side is where Google points customers with document workloads: parsing documents, reading charts and data, and drafting reports off them. It ships with strengthened safeguards in the chemical, biological, radiological, nuclear and cyber-offence domains, which Google says make it substantially harder to jailbreak while being trained to refuse less on legitimate requests.

The honest caveat is the calendar. Gemini 3.7 Flash arrived three weeks later at an introductory price Google set at half of 3.6 Flash's original rate, and the lab's own comparison puts it ahead of 3.6 Flash on every coding, document and workflow row it publishes. 3.6 Flash remains a reasonable choice where a workload has already been tuned against it, not where a new one is being started.

What Google announced at launch

Cost per task, not per token
Google reports fewer output tokens than 3.5 Flash for better quality, and fewer reasoning steps and tool calls per multi-step workflow, at a lower price per token than the model it replaced.
Cleaner code edits
On DeepSWE the lab reports 49% against 37% for 3.5 Flash, attributing the gain to higher precision, fewer unwanted edits and fewer execution loops rather than to more attempts.
Research and computer use
Google reports 63.9% against 49.7% on MLE Bench for machine learning research work, and 83.0% against 78.4% on OSWorld-Verified, with computer use now a built-in client-side tool.
Document and chart work
The lab points customers with document workloads here: parsing documents, analysing charts and data, and drafting reports from them, with knowledge work measured at 1421 Elo against 1349.
Jailbreak resistance
It ships with strengthened CBRN and cyber-offence safeguards that Google says make it substantially more resistant to jailbreaks, while being trained to refuse fewer beneficial requests.
What it is not for
Gemini 3.7 Flash followed three weeks later at the same introductory rate, and Google's own figures put it ahead on every coding, document and workflow row. New work belongs on the newer model.

Indian languages

Gemini 3.6 Flash answers in 15 Indian languages. Choose a language from the menu beside the message box to receive replies in it.

Frequently Asked Questions

Frequently asked questions about Gemini 3.6 Flash.

When was Gemini 3.6 Flash released?

Google released Gemini 3.6 Flash on Jul 21, 2026.

Who built Gemini 3.6 Flash?

Gemini 3.6 Flash is developed by Google. 99Models AI connects directly to it at the provider's published rate.

How intelligent is Gemini 3.6 Flash?

It is ranked 28 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does Gemini 3.6 Flash cost?

Usage costs ₹79.20 per million input tokens and ₹396.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is Gemini 3.6 Flash pricing in US dollars?

The provider charges $0.83 per million input tokens and $4.13 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can Gemini 3.6 Flash hold?

Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does Gemini 3.6 Flash rank for value?

It ranks 17 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does Gemini 3.6 Flash support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

Which Indian languages does Gemini 3.6 Flash support?

It answers in 15 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.

More models from Google

Catalog updated Sep 9, 2026