DeepSeek officially launched V4-Pro-0813 on August 13, 2026, available via API, app, and web, including via 'Expert Mode' on the app and web. The model is a Mixture-of-Experts architecture with 1.6 trillion total parameters and 49 billion activated parameters, compared to predecessor DeepSeek-V3.2-Base's 671 billion total and 37 billion activated. DeepSeek published open weights on Hugging Face under the MIT license alongside a technical report. The company said V4-Pro greatly enhances agent capabilities. The release includes native OpenAI Responses API support with one-click Codex integration and thinking effort in three tiers: low, high, and max. API model names are unchanged from prior versions.
On benchmarks, V4-Pro scores 93.5 on LiveCodeBench Pass@1, 80.6 on SWE Verified, and 83.5% on MRCR 1M needle-in-a-haystack retrieval, surpassing Gemini 3.1 Pro on that last test. Its Hybrid Attention mechanism requires only 10% of the KV cache that V3.2 needed at 1M context. On Code Arena WebDev AutoEval it reached 1607 points, placing it around eighth overall and second among open models, 15 Arena points behind GPT-5.6 Sol xHigh while costing approximately 1/31st of that model's blended token price. Kimi K3 Max leads at 1674, above GPT-5.6 Sol at 1622; for a 1M-input plus 1M-output workload, Kimi costs $18 versus DeepSeek's $1.305. Those Code Arena scores are early AutoEval results in which a reward model casts automatic votes in place of live human votes, so the practical quality gap in real coding tasks may differ from the point difference. GPT-5.5 leads on Terminal-Bench 2.0 at 82.7% versus V4-Pro's 67.9% and on GPQA Diamond at 93.6% versus 90.1%. DeepSeek claims V4-Pro trails the absolute frontier by roughly 3 to 6 months. MindStudio reported V4-Pro at approximately 91.2% on SWE-Bench Verified versus Claude Opus 4.7's 93.9%. Kilo.ai's FlowGraph coding test placed V4-Pro at 77/100, between Claude Opus 4.7 at 91 and Kimi K2.6 at 68; V4 Flash scored 60/100 at $0.02 per run but its build failed with missing key components.
Before the official release, the model was accessible via API on OpenRouter. Simon Willison observed that V4-Pro produces noticeably different outputs across its low, medium, and high reasoning levels, writing he had 'not noticed this kind of difference from any other model.' Benchmark results reportedly spread from DeepSeek's WeChat group to a deleted Reddit post to an ASCII table on Hacker News.
API output token pricing at peak hours rose to $3.96 per million, up from a prior flat rate of $0.87 per million, with peak and off-peak tiers introduced. V4 Pro prices are up to 14 times higher than the V4 Flash tier. DeepSeek had offered a 75% promotional discount on V4-Pro through May 5. DeepSeek's rates remain below Anthropic's Fable 5 at $50 per million output tokens. DeepSeek said the release is part of broader expansion including hiring, computing capacity, and fundraising.