DeepSeek V4.1 Flash
NewMultimodal Flash successor for efficient reasoning and agent workflows.
Released Sep 10, 2026
₹144.75
per 10 lakh output tokens
Input: ₹36.19 per 10 lakh tokens
Billed at provider rates converted at ₹96.5 per US dollar, with 0% markup.
Pricing
Benchmarks
Scores are percentages unless marked as a rating. All benchmarks are measured independently.
89.1%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
39.2%
HLE
Humanity's Last Exam
51.9%
SciCode
SciCode - scientific code generation
84.0%
Long Context
Long Context Reasoning - reasoning over long inputs
About DeepSeek V4.1 Flash
A multimodal Flash successor for reasoning, coding and agent workflows, with native image input and a million-token context.
V4.1 Flash introduces DeepSeek's Causal Encoder-Decoder architecture and native visual understanding. It is intended for reasoning and agent workflows where throughput and the cost of repeatedly reading context matter alongside answer quality.
The design assigns different amounts of computation to reading and writing: eight billion active parameters for input and sixteen billion for output. DeepSeek attributes its capability gains to new pretraining methods and larger-scale reinforcement learning, and reports results ahead of V4 Pro on its evaluations.
The release also reduces the memory and storage required by the KV cache. DeepSeek reports one quarter of the previous generation's HBM footprint and one eighth of its SSD requirement. That efficiency is particularly relevant to agents that revisit long contexts across many steps; actual serving limits and billing depend on the selected host.
What DeepSeek announced at launch
- Asymmetric computation
- Input and output use different active parameter budgets within a 552B MoE.
- Native vision
- Visual understanding is integrated into the new Flash architecture.
- Smaller context cache
- DeepSeek reports substantial reductions in cache memory and storage needs.
Indian languages
DeepSeek V4.1 Flash answers in 12 Indian languages. Choose a language from the menu beside the message box to receive replies in it.
Frequently Asked Questions
Frequently asked questions about DeepSeek V4.1 Flash.
When was DeepSeek V4.1 Flash released?
DeepSeek released DeepSeek V4.1 Flash on Sep 10, 2026.
Who built DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash is developed by DeepSeek. 99Models AI connects directly to it at the provider's published rate.
How intelligent is DeepSeek V4.1 Flash?
It is ranked 19 of 56 chat models on intelligence, ordered by independent benchmark scores. View its complete scores in the Benchmarks table above.
How much does DeepSeek V4.1 Flash cost?
Usage costs ₹36.19 per million input tokens and ₹144.75 per million output tokens, with 0% markup. There is no subscription; you pay only for what you use.
What is DeepSeek V4.1 Flash pricing in US dollars?
The provider charges $0.38 per million input tokens and $1.50 per million output tokens. Rupee rates are converted at ₹96.5 per US dollar.
How long a conversation can DeepSeek V4.1 Flash hold?
Its context window is 10.5 lakh tokens. That is the total volume of text and attached files it can process in a single request.
How does DeepSeek V4.1 Flash rank for value?
It ranks 5 out of 56 models for value. This ranking weighs benchmark intelligence against the actual token cost.
Does DeepSeek V4.1 Flash support reasoning?
On by default. Where reasoning is supported, you can adjust the thinking effort level directly in the message composer.
Which Indian languages does DeepSeek V4.1 Flash support?
It answers in 12 Indian languages. Select your preferred language from the menu beside the message box to receive replies in it.
More models from DeepSeek
Catalog updated Sep 20, 2026