Key Takeaways
- DeepSeek V4.1 Flash is currently 1,500 times cheaper than usual on some third-party hosts.
- Input pricing on Relace and OpenRouter is as low as $0.0001 per million tokens.
- Output pricing remains close to DeepSeek’s off-peak rates, limiting overall savings.
DeepSeek V4.1 Flash, a recently launched model, has seen its input pricing plummet by 1,500 times on some third-party hosts, according to a Reddit thread on r/DeepSeek.
On platforms like OpenRouter, Relace lists the model at $0.0001 per million input tokens, while DeepSeek’s own off-peak input price is $0.15 per million tokens.
While the input cost is significantly reduced, the output pricing on Relace still matches DeepSeek’s off-peak rate of $0.60 per million tokens, meaning developers may not see the same savings for chatty or long-answer jobs.
Open Inference’s OpenRouter listing shows a slightly higher input price of about $0.00011 per million tokens and $0.36 per million output tokens, but recent snapshots have shown weaker uptime and slower throughput compared to Relace.
The DeepSeek V4.1 Flash model, released around September 10, 2026, is a sparse mixture-of-experts model built on the company’s Causal Encoder-Decoder architecture, with a 552 billion parameter backbone.
The model is designed for coding, terminal work, computer-use agents, and long-context jobs, with a 1 million token context window and compressed KV caching to reduce memory usage.
Despite the significant price drop, developers are advised to weigh the cost against reliability and speed when choosing a host, as output quality and service stability can vary.
This price drop comes in the wake of earlier pricing moves by DeepSeek, including a 4x increase in prices, highlighting the volatility in the market.
The current listings on Relace and OpenRouter represent a stark contrast to DeepSeek’s official pricing, sparking interest among developers for high-volume input workloads.





