DeepSeek V4 Preview Review: The Ultimate Breakdown of Specs, Tests, and Real-World Performance
The DeepSeek team is back with their highly anticipated V4 Preview, built around a massive 1 million context length. Released under the MIT license — a huge win for the open-source community — the release includes two distinct models: the flagship Pro and the cost-effective Flash.
DeepSeek claims these models are the top open-source performers, rivaling closed-source giants and excelling in reasoning, STEM, coding, and agentic workflows. But does the real-world performance live up to the impressive benchmarks? Let's dive deep into the specs, pricing, and extensive testing.
Under the Hood: Specs and Pricing
DeepSeek offers two distinct flavors in this preview release, both highly efficient and incredibly cheap.
DeepSeek V4 Pro
The Flagship · model id: deepseek-v4-pro
Size: 1.6T total params · 49B active
Context: 1M tokens · Max output 384K
Input (cache miss): $0.435 / 1M tokens
Input (cache hit): $0.003625 / 1M tokens
Output: $0.87 / 1M tokens
DeepSeek V4 Flash
The Lightweight Alternative · model id: deepseek-v4-flash
Size: 284B total params · 13B active
Context: 1M tokens · Max output 384K
Input (cache miss): $0.14 / 1M tokens
Input (cache hit): $0.0028 / 1M tokens
Output: $0.28 / 1M tokens
Rates are per 1M tokens in USD, read from our single pricing source of truth (last verified 2026-07-25). See DeepSeek pricing for the full rate card and price history.
Availability: Open weights are available on HuggingFace and Ollama Cloud, or free via the official DeepSeek chatbot. The API is OpenAI- and Anthropic-compatible (https://api.deepseek.com · /anthropic) and supports JSON mode, tool calls, FIM completion, and chat-prefix completion.
Heads up — deprecation: the legacy deepseek-chat and deepseek-reasoner model ids will be deprecated on 2026-07-24. They map to V4 Flash's non-thinking and thinking modes respectively, so plan a migration to deepseek-v4-flash or deepseek-v4-pro well before that date.
Benchmarks vs. Reality: Is it "Benchmark Maxed"?
DeepSeek positioned V4 against Gemini 3.1 Pro and GPT-5.4 on several reasoning and coding benchmarks. It made no claim against Claude Opus 4.8 — Opus 4.8 shipped after V4, so any "V4 beats Opus 4.8" line you see online is a third-party comparison, not a DeepSeek claim. Hands-on, the models read as partly "benchmark maxed": strong on standard tests, less consistent in practical use.
Community leaderboards back that up. On Code Arena, deepseek-v4-prosits at #35 and its thinking variant at #31 — well behind other Chinese models such as GLM 5.1 and Kimi K2.6, and behind the Western frontier. Leaderboard positions move constantly; check the live board before quoting them.
The Backlash and the Verdict
Calling the DeepSeek V4 models "mid" has sparked significant backlash on Twitter, particularly from the Chinese AI community. The honest reading is narrower: on public coding leaderboards the V4 family currently ranks below GLM 5.1 and Kimi K2.6, and below the Western frontier including Claude Opus 4.8 — which launched after V4 and which DeepSeek never benchmarked against. What V4 does win on is price per token, by a wide margin.
Pros
• Massive 1M context length
• MIT license — fully open
• Incredibly cheap pricing
• Flash model punches above its weight
• Strong on 360° / 3D rotational tasks
Cons
• Feels "benchmark maxed"
• Sloppy real-world execution
• Lags behind GLM 5.1 & Qwen 3.6 Plus
• Front-end output looks dated
• Pro model fails complex agentic flows
While the pricing, context length, and efficiency make DeepSeek V4 an incredible base for future development, this preview version needs serious refinement. At the end of the day, cheaper doesn't make it better — it just means it's cheaper.
Try DeepSeek in your browser
Access DeepSeek V4 and other models directly from any webpage with the AI Sidebar Chrome Extension.
Install AI Sidebar