99MODELS

Gemini 3.8 Flash

New

Latest Flash generation; fast multimodal reasoning across every input type.

Released Sep 2, 2026

360.00

per 10 lakh output tokens

Input: ₹72.00 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
10,48,576 tokens
Max output
65,536 tokens
Accepts
Text, Images, PDF files, Audio, Video
Reasoning
On by default
Effort levels
low, medium, high
Tool use
Yes
Structured output
Yes
Code execution
Yes
Intelligence rank
#12 of 54
Value rank
#7 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input72.00$0.75
Output360.00$3.75
Cached input7.20$0.08

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 95.3%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 47.8%

    HLE

    Humanity's Last Exam

  • 56.6%

    SciCode

    SciCode - scientific code generation

  • 81.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 87.6%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

  • 1567

    WebDev Arena

    WebDev Arena - head-to-head web-app builds, Elo rating

About Gemini 3.8 Flash

Google's latest workhorse Flash model, delivering its predecessor's speed and price with a marked step up in software engineering, agentic work and multi-step reasoning in specialised domains. It accepts text, images, audio, video and files across a million-token context, and carries three thinking levels. Google is explicit that it works harder than 3.7 Flash -- more reasoning steps and more iterative tool calls -- so it can spend more tokens on the same task, especially at higher effort.

This is the third Flash release in six weeks, and the pattern holding across all three is that the price does not move. Google calls 3.8 Flash its most intelligent workhorse model and puts it out at exactly what 3.7 Flash costs, aiming the gains at software engineering, agentic tasks and multi-step reasoning in specialised domains rather than at general chat.

Its reported results follow that aim. On DeepSWE v1.1, which measures long-horizon software engineering, Google says it outperforms most larger frontier models at autonomously solving complex engineering problems end to end, at a fraction of their cost. It reports 54.9% on HLE-Verified for multi-step reasoning across STEM, humanities and professional fields, and says it exceeds 3.7 Flash and other frontier models on the Vals Finance Agent V2 and Harvey Legal Agent benchmarks - two evaluations pointed at professional domains rather than at coding. Google also reports a significant leap in prompt-injection robustness on Gray Swan, which matters more than usual for a model this cheap to run in an agent loop.

The interesting caveat is one Google raises itself, and it is a cost caveat rather than a quality one. It says 3.8 Flash "works harder": it executes extra reasoning steps and calls tools iteratively, and it may therefore use more tokens to reach a better answer, especially at higher effort. Google's own recommendation is that if efficiency matters more than the last few points of quality, 3.7 Flash remains the model to reach for - the two are siblings on the same price sheet, not a replacement and a legacy.

The price sheet has a date on it, which is the one thing about this model most likely to catch someone out. The $0.75 input and $3.75 output rates run through 31 December 2026 and double to $1.50 and $7.50 on 1 January 2027, with context caching moving from $0.075 to $0.15 in step. That is introductory pricing, not a permanent rate.

Google shipped a second model in the same post, Gemini 3.8 Flash Cyber, its most capable cybersecurity model, aimed at vulnerability detection and automated patching and available to trusted defenders through a new Fairwind Program. That one is not served here; everything above describes the general model.

What Google announced at launch

Same price, more capability
Google's third Flash release in six weeks holds 3.7 Flash's rates while improving software engineering, agentic work and multi-step reasoning in specialised domains.
Long-horizon software engineering
On DeepSWE v1.1 Google reports it outperforming most larger frontier models at solving complex engineering problems end to end autonomously, at a fraction of their cost.
Reasoning outside code
It reports 54.9% on HLE-Verified across STEM, humanities and professional fields, and says it exceeds 3.7 Flash and other frontier models on the Vals Finance Agent V2 and Harvey Legal Agent benchmarks.
Harder to inject
Google reports a significant leap in prompt-injection robustness on Gray Swan - the property that decides whether a cheap model is safe to leave running in an agent loop.
Introductory pricing, with a date
The $0.75 and $3.75 rates run through 31 December 2026 and double on 1 January 2027, with cached input moving from $0.075 to $0.15 at the same time.
What it is not for
Google says this model works harder - more reasoning steps, more iterative tool calls - and may spend more tokens for the same task at higher effort. It points at 3.7 Flash when efficiency matters more than the last few points.

Indian languages

Gemini 3.8 Flash answers in 15 Indian languages. Choose a language from the menu beside the message box to receive replies in it.

Frequently Asked Questions

Frequently asked questions about Gemini 3.8 Flash.

When was Gemini 3.8 Flash released?

Google released Gemini 3.8 Flash on Sep 2, 2026.

Who built Gemini 3.8 Flash?

Gemini 3.8 Flash is developed by Google. 99Models AI connects directly to it at the provider's published rate.

How intelligent is Gemini 3.8 Flash?

It is ranked 12 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does Gemini 3.8 Flash cost?

Usage costs ₹72.00 per million input tokens and ₹360.00 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is Gemini 3.8 Flash pricing in US dollars?

The provider charges $0.75 per million input tokens and $3.75 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can Gemini 3.8 Flash hold?

Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does Gemini 3.8 Flash rank for value?

It ranks 7 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does Gemini 3.8 Flash support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

Which Indian languages does Gemini 3.8 Flash support?

It answers in 15 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.

More models from Google

Catalog updated Sep 9, 2026