99MODELS

DeepSeek V4 Flash Vision

New

Experimental vision build of V4 Flash; reads images at V4 Flash prices.

Released Aug 21, 2026

126.72

per 10 lakh output tokens

Input: ₹42.24 per 10 lakh tokens

Billed at provider rates converted at ₹96 per US dollar, with 0% markup.

Specifications

Context window
10,48,576 tokens
Max output
2,62,144 tokens
Accepts
Text, Images
Reasoning
On by default
Effort levels
low, high, max
Tool use
Yes
Structured output
Yes
Code execution
No
Intelligence rank
#27 of 54
Value rank
#6 of 54

Pricing

Pricing
Per 10 lakh tokensINRUSD
Input42.24$0.44
Output126.72$1.32
Cached input13.44$0.14

Benchmarks

Scores are percentages unless marked as a rating. All benchmarks are measured independently.

  • 91.3%

    GPQA Diamond

    GPQA Diamond - graduate-level science Q&A

  • 34.5%

    HLE

    Humanity's Last Exam

  • 49.7%

    SciCode

    SciCode - scientific code generation

  • 81.3%

    Long Context

    Long Context Reasoning - reasoning over long inputs

  • 74.2%

    Terminal-Bench 2

    Terminal-Bench 2.1 - agentic terminal tasks, second edition

About DeepSeek V4 Flash Vision

An experimental vision build of V4 Flash: the same sparse Mixture-of-Experts design with 13B active parameters out of 284B total, extended to read images while matching the base model on text, agents, reasoning and world knowledge. DeepSeek aims it at document and chart understanding, visual question answering and multimodal agent workflows that interleave text and images, and reports its multimodal agent scores as a large jump over plain V4 Flash. Images are tokenised for billing at up to 384 tokens each, so vision costs V4 Flash rates. It is labelled experimental by DeepSeek, and the id may change.

This is V4 Flash with eyes, and DeepSeek is careful to say it is nothing more than that on the text side. It is the same sparse Mixture-of-Experts design - 13B active parameters out of 284B total - and the lab's claim is that it matches plain V4 Flash on text capabilities including agents, reasoning and world knowledge, with image understanding added rather than traded for.

Where it moves is the multimodal half. DeepSeek reports that on multimodal agent benchmarks the model makes a major leap over V4 Flash, bringing multimodal agent performance close to Claude Opus 4.8. That is the claim the release exists to make, and it is a specific one: not that the model sees well in isolation, but that it holds up when seeing is one step inside a longer agentic loop.

The uses DeepSeek names follow from that framing - document and chart understanding, visual question answering, and multimodal agent workflows that interleave text and images. It is the interleaving that matters. A model that reads a chart is useful; a model that reads a chart in the middle of a tool-calling sequence, without losing the thread of what it was doing, is a different capability, and it is the one being claimed here.

Billing is the part most likely to surprise. Images are tokenised at up to 384 tokens each and charged at ordinary V4 Flash rates, so vision costs essentially nothing beyond the tokens - a screenshot is priced like a short paragraph. Access to the Files API is free. Against a model already among the cheapest capable options available, that makes visual input unusually inexpensive to use in volume.

The caveat is in the name and DeepSeek does not soften it: this is labelled experimental, and its id carries an "exp" suffix to say so. Experimental ids get renamed, superseded or withdrawn on the lab's schedule rather than yours, so treat it as a capability to try rather than one to build a dependency on. There is no separate published architecture or context detail beyond what V4 Flash already carries.

What DeepSeek announced at launch

V4 Flash, plus sight
The same sparse Mixture-of-Experts model with 13B active parameters out of 284B total. DeepSeek says it matches plain V4 Flash on text, agents, reasoning and world knowledge, with vision added rather than substituted.
Multimodal agents, not just images
DeepSeek reports a major leap over V4 Flash on multimodal agent benchmarks, bringing that performance close to Claude Opus 4.8 - the claim is about seeing inside a longer loop, not about single-image tasks.
What it is built to read
Documents and charts, visual question answering, and agent workflows that interleave text and images without losing the thread of the task in between.
Vision at text prices
Images are tokenised at up to 384 tokens each and billed at ordinary V4 Flash rates, so a screenshot costs about what a short paragraph costs. Files API access is free.
What it is not for
DeepSeek labels this experimental. Experimental ids get renamed, superseded or withdrawn on the lab's timetable, so it is a capability worth trying rather than one to build a hard dependency on.

Frequently Asked Questions

Frequently asked questions about DeepSeek V4 Flash Vision.

When was DeepSeek V4 Flash Vision released?

DeepSeek released DeepSeek V4 Flash Vision on Aug 21, 2026.

Who built DeepSeek V4 Flash Vision?

DeepSeek V4 Flash Vision is developed by DeepSeek. 99Models AI connects directly to it at the provider's published rate.

How intelligent is DeepSeek V4 Flash Vision?

It is ranked 27 of 54 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.

How much does DeepSeek V4 Flash Vision cost?

Usage costs ₹42.24 per million input tokens and ₹126.72 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.

What is DeepSeek V4 Flash Vision pricing in US dollars?

The provider charges $0.44 per million input tokens and $1.32 per million output tokens. Rupee rates are converted at ₹96 per US dollar.

How long a conversation can DeepSeek V4 Flash Vision hold?

Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.

How does DeepSeek V4 Flash Vision rank for value?

It ranks 6 out of 54 models for value. This ranking weighs benchmark intelligence against the actual token cost.

Does DeepSeek V4 Flash Vision support reasoning?

On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.

More models from DeepSeek

Catalog updated Sep 9, 2026