Key Takeaways
- DeepSeek has launched V4.1 Flash, replacing older models with improved features.
- The new model offers lower API prices and native image understanding capabilities.
- V4.1 Flash supports up to 1 million tokens and has a context window of 384,000 tokens.
DeepSeek has officially launched V4.1 Flash, a significant update to its existing Flash models. The new version, which became available on September 10, 2026, under the API name deepseek-flash, introduces several enhancements and cost reductions.
According to the company, V4.1 Flash uses a redesigned architecture with approximately 748 billion total parameters, including a 552 billion backbone and 196 billion Engram parameters. Despite its size, the model operates efficiently, with only around 8 billion parameters active during prefill and 16 billion during decoding, helping to keep inference costs lower.
One of the key features of V4.1 Flash is its native multimodal capabilities. The DeepSeek-ViT encoder can process images alongside text, supporting resolutions up to roughly 1344 × 1344 pixels. This feature allows for more versatile and integrated data processing, making it suitable for a wide range of applications.
The model supports a context window of up to 1 million tokens and can generate outputs of up to 384,000 tokens. Developers can also adjust the reasoning effort from 1 to 100, providing a balance between cost and accuracy.
DeepSeek’s own benchmarks show significant improvements over its previous models. The DeepSWE score increased from 54.4 on V4 Flash to 74.2, while Terminal Bench 2.1 rose from 82.7 to 90.6. V4.1 Flash also scored 88.1 on CyberGym and 54.8 on AutomationBench. Independent testing has also been positive, with Artificial Analysis reporting a 68.9% score on AutomationBench-AA and Vals.ai giving the model a 57.86% Vals Index.
In terms of pricing, DeepSeek has reduced API costs significantly. During off-peak hours, cached input now costs $0.003 per million tokens, uncached input costs $0.15, and output costs $0.60. Peak pricing doubles these figures to $0.006 for cached input, $0.30 for uncached input, and $1.20 for output per million tokens. The biggest reduction applies to cached input, which should particularly benefit agent workloads that repeatedly reuse large context windows.
DeepSeek is also preparing to phase out V4 Pro. Starting September 14 at 04:00 UTC, requests sent to deepseek-v4-pro will automatically be redirected to V4.1 Flash and charged at the new Flash rates. This arrangement will remain in place until V4.1 Pro becomes available, potentially reducing costs for existing V4 Pro users.
While V4.1 Flash offers substantial improvements and cost savings, it does not lead every benchmark. It trails some frontier models on newer Terminal Bench tests and performs poorly on certain specialized evaluations such as SRE Bench and Harvey’s Legal Agent benchmark.





