DeepSeek Retires Flagship V4-Pro in Favour of Cheaper Flash Model
Chinese AI lab DeepSeek launched V4.1-Flash on Sept 10 and will reroute all requests to its former flagship V4-Pro to the new model from Sept 14, billed at Flash rates. The company says its smaller, cheaper model outperforms the 1.6-trillion-parameter Pro on coding and agent tasks while cutting costs.
Bottom line — DeepSeek's V4.1-Flash activates 8-16B parameters per token and costs €0.55 per million output tokens off-peak.
Go deeper (8)
- DeepSeek reports V4.1-Flash uses 552B total parameters but activates only 8B during input and 16B during output, a design aimed at agent workloads that consume large amounts of context.
- On DeepSeek's own benchmarks, V4.1-Flash scored 74.2 on DeepSWE v1.1, edging Claude Opus 5 (74.0) and GPT-5.6 Sol (73.0), according to the company's figures.
- The model lags on reasoning: on Humanity's Last Exam it scored 36.8 versus Opus 5's 56.3, per DeepSeek's data.
- Bloomberg Intelligence analysts estimate the price cut at up to 32%; V4.1-Flash costs €0.55 per million output tokens off-peak, with cached input at €0.003 (DeepSeek's rates).
- Competitor shares fell: MiniMax and Z.ai dropped more than 8% in Hong Kong trading on Thursday, per Bloomberg's Saritha Rai.
- The model is released under MIT license on Hugging Face, but its 510GB size makes local deployment challenging, with community builds only running on machines with at least 512GB RAM.
- DeepSeek's API changelog initially stated V4-Pro would be retired, but a later update said the service would continue—both pages were live as of Sept 11, per Frontier Watch.
- A separate earlier deprecation: DeepSeek's legacy model names deepseek-chat and deepseek-reasoner were shut down on July 24, 2026, with no compatibility shim (Times Tabloid).