Key Takeaways
- DeepSeek's V4-Flash model outperforms GLM 5.2 in benchmark tests.
- It costs significantly less than Claude Fable and other models.
- V4-Flash supports up to 2,500 concurrent requests per account.
DeepSeek has released its latest lightweight model, V4 Flash, which outperforms top models such as GLM 5.2 in several benchmark tests while costing a fraction of the price.
According to DeepSeek's reports, V4-Flash achieved higher scores than GLM-5.2 across nine different benchmarks, including Terminal Bench 2.1 and NL2Repo.
The model is also competitive with Claude Opus 4.8 in certain tests, coming close on Terminal Bench 2.1, DeepSWE, Agents’ Last Exam, AutomationBench Public, DSBench-FullStack, and DSBench-Hard.
DeepSeek charges $0.14 per million uncached input tokens, $0.0028 per million cached input tokens, and $0.28 per million output tokens for V4-Flash, making it significantly cheaper than V4-Pro at $0.435 per million uncached input tokens and $0.87 per million output tokens.
The API supports up to 2,500 concurrent requests per account, compared with 500 for V4-Pro, and includes features such as thinking and non-thinking modes, a one-million-token context window, and outputs of up to 384,000 tokens.
DeepSeek has open-sourced the model under the MIT licence, allowing commercial and on-premise deployment. The underlying mixture-of-experts model has 284 billion parameters with 13 billion activated per token, supporting a one-million-token context window.
The V4-Flash API now natively supports the Responses API format and has been adapted for Codex, enhancing its utility in various applications.
DeepSeek did not update the V4-Pro API or models used in its web and mobile applications. The official V4-Pro model is still under development with no announced release date.





