Short answer: on list price, yes — but read the harness caveats. On July 31, 2026, DeepSeek promoted V4-Flash-0731 to an official API beta: same 284B total / 13B active architecture as April's preview, with gains from post-training only. It beats DeepSeek's own larger V4-Pro preview on agent benchmarks at roughly 1/36 to 1/179 of Claude Opus 4.8's price. For developers evaluating frontier model selection and Agent workflow costs, this guide covers: the April–August release timeline, official pricing tables, CSA+HCA (DSA) architecture, the first mention of DeepSeek Harness, six-way comparison with Artificial Analysis Intelligence Index and per-task cost, controversies (harness dependency, cache-hit issues, unconfirmed V4-Pro GA), the Chinese "kill line" concept and Liang Wenfeng nickname saga, six operational steps, and FAQ. Key hard data: Flash input $0.14 / $0.0028 (cache miss/hit), output $0.28 per million tokens; V4-Pro at 1M context needs 27% of V3.2 FLOPs and 10% KV cache; Terminal Bench 2.0 82.7 vs V4-Pro preview 67.9 — all agent scores run on unreleased Harness minimal mode.
DeepSeek's V4 rollout was a three-month sequence, not a single launch day. Each step shifted competitive pressure against Kimi K3, Qwen3.8-Max, and Western closed models.
| Date | Event |
|---|---|
| April 24, 2026 | DeepSeek-V4 preview launches and open-sources two MIT-licensed MoE models: V4-Pro (1.6T total / 49B active) and V4-Flash (284B total / 13B active), both with 1M-token context |
| July 24, 2026 | Legacy aliases deepseek-chat and deepseek-reasoner are retired; all traffic routes to V4 family naming |
| July 27, 2026 | Moonshot AI ships Kimi K3 open weights on Hugging Face (2.8T total parameters) — one of the largest open-weight releases to date, raising pressure on DeepSeek days before its own update |
| July 31, 2026 | V4-Flash-0731 official API beta goes live; open weights sync to Hugging Face under MIT. Changelog names "DeepSeek Harness" for the first time, calling it "to be released soon." API-only — consumer app and web chat not updated |
| As of August 5, 2026 | V4-Pro official GA remains unconfirmed. Chinese outlets citing unnamed sources report internal testing began the week of July 28 with a possible GA window of August 10–20 — not confirmed by DeepSeek; treat as rumor |
Verification note: Data in this article is current as of August 5, 2026. V4-Pro GA dates, Harness public release, peak-hour surcharge timing, and third-party benchmark reproductions may change — confirm against DeepSeek's official changelog and API docs before migrating production workloads.
| Model | Status | Total / active params | Context | Input (cache miss / hit, per 1M tokens) | Output (per 1M tokens) | License |
|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash-0731 | Official (Jul 31, 2026) | 284B / 13B | 1M | $0.14 / $0.0028 | $0.28 | MIT |
| DeepSeek-V4-Pro | Preview (Apr 24, 2026) | 1.6T / 49B | 1M | $0.435 / $0.003625 | $0.87 | MIT |
| Kimi K3 (Moonshot AI) | Open weights (Jul 27, 2026) | 2.8T / ~104B (community estimate) | ~1.05M | $3.00 / $0.30 | $15.00 | Modified MIT |
| GLM-5.2 (Zhipu / Z.ai) | Open (June 2026) | ~744B / ~40B | 1M | Not verified for this piece | Not verified | MIT |
| Qwen3.8-Max (Alibaba) | API GA (Aug 2, 2026); weights pending | 2.4T / 95B | 1M | $2.00 / ~$0.17–0.25 | $6.00 | Open weights promised |
All pricing figures are vendor-published rates. DeepSeek has announced a future 2x peak-hour surcharge (9am–12pm and 2pm–6pm Beijing time) with no confirmed effective date yet.
The most easily missed detail: V4-Flash-0731 is identical in size and structure to April's preview. DeepSeek states the entire performance jump on agent benchmarks came from re-running post-training, not scaling up. A 284B/13B model now beats a 1.6T/49B sibling on multiple agentic tasks — cutting against the default assumption that bigger equals better.
DeepSeek's technical report (DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence) describes three architectural changes carried from the April preview:
At 1M-token context, DeepSeek claims V4-Pro needs only 27% of V3.2's per-token inference FLOPs and 10% of KV cache footprint. These are vendor-reported efficiency numbers; independent third-party reproduction of these specific figures has not been published as of August 5, 2026.
July 31's changelog marked the first official mention of DeepSeek Harness — an in-house agent execution framework for file I/O, tool calls, and multi-step engineering tasks, positioned as DeepSeek's answer to Claude Code. Until now, DeepSeek's teams relied on Claude Code, OpenCode, and other third-party agent tools.
Every agent benchmark DeepSeek published for V4-Flash-0731 (Terminal Bench 2.0, Toolathlon, etc.) was measured using Harness's "minimal mode," which is not yet publicly released, at max reasoning effort, top_p 0.95, temperature 1.0. DeepSeek's changelog adds that agent scores are "extremely sensitive to harness choice" — a caveat worth taking at face value rather than skipping past.
Headline agent scores are harness-dependent and self-reported. Until third parties reproduce with Claude Code, Cursor, or other frameworks, treat Terminal Bench 2.0 numbers as "vendor plus specific framework" results, not portable capability claims.
V4-Flash-0731 lands in the densest stretch of Chinese open-weight model launches this year. Artificial Analysis Intelligence Index and per-task cost figures below are independent; DeepSeek's own agent benchmarks use a different methodology and are listed separately in the source material.
| Model | Lab | Release / weights | Total params | AA Intelligence Index | Avg. cost per task (AA) |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash-0731 | DeepSeek | Jul 31, 2026 (official) | 284B | 50 | $0.03 |
| Kimi K3 | Moonshot AI | Jul 16 preview / Jul 27 weights | 2.8T | 57 | $0.86 |
| GLM-5.2 | Zhipu / Z.ai | June 2026 | ~744B | ~1 pt above V4-Flash | Not verified |
| Qwen3.8-Max | Alibaba | Aug 2, 2026 GA (weights pending) | 2.4T | Not verified | Not verified |
| GPT-5.6 Sol | OpenAI | Closed source | Undisclosed | 9+ pts above V4-Flash | $1.86 |
| Claude Fable 5 | Anthropic | Closed source | Undisclosed | 9+ pts above V4-Flash | $3.15 |
Sourcing: Intelligence Index and per-task cost from Artificial Analysis, as reported by financial outlets (Wantrich, Meyka). DeepSeek agent benchmarks (Terminal Bench 2.0, SWE-bench Verified, etc.) are vendor-reported via unreleased Harness minimal mode — do not conflate the two datasets.
The tension: on Artificial Analysis's independent index, V4-Flash trails Kimi K3 and GLM-5.2. But per-task cost is roughly 1/29th of Kimi K3, 1/62nd of GPT-5.6 Sol, and 1/105th of Claude Fable 5. DeepSeek is not competing for leaderboard top spot — it is optimizing for "good enough intelligence at a price nobody else can match," which is also why the preview reportedly topped OpenRouter's most-used model ranking for seven consecutive weeks.
Before V4-Flash-0731 shipped, Chinese AI forums mocked DeepSeek founder Liang Wenfeng as "Liang Baikai" — a pun roughly meaning "Liang Empty Promise," after V4-Pro's mid-July target slipped. Once the official Flash build outperformed expectations, communities flipped back to "Liang Sheng" ("Liang the Sage") — a small but telling barometer of sentiment swings in China's AI developer community.
More substantively, Chinese developer circles use "斩杀线" (zhǎn shā xiàn — "kill line"): DeepSeek's combination of good-enough performance plus rock-bottom price sets an effective bar. Competitors whose models don't clearly beat DeepSeek on capability and can't undercut it on price risk losing market relevance. That framing helps explain moves like OpenAI reportedly cutting GPT-5.6 Luna prices by 80% around the same period. One AI incubator source quoted by 21st Century Business Herald: "Every large-model company is running ahead of Liang Wenfeng — however they do it, they have to stay ahead of DeepSeek to survive."
On July 31, Nvidia, Broadcom, and AMD saw no significant stock movement when V4-Flash went official — a contrast to early 2025, when DeepSeek-R1's efficiency claims triggered a global AI-chip selloff. Markets appear to treat "DeepSeek does more with less compute" as normal engineering rather than an automatic bearish signal for compute demand.
Verify your API endpoint: confirm deepseek-v4-flash resolves to the 0731 official build; legacy deepseek-chat and deepseek-reasoner were retired July 24
Choose Flash vs Pro for your workload: V4-Flash-0731 wins on cost and beats V4-Pro preview on Harness-reported agent scores; use V4-Pro preview only for deeper world-knowledge tasks until official GA
Model cost with cache tiers: bill at $0.14 cache-miss vs $0.0028 cache-hit input — cache strategy materially changes the Claude price-gap math (36x to 179x on input)
Treat Harness benchmarks as framework-specific: do not port Terminal Bench 2.0 scores to Claude Code or Cursor without independent reproduction on your stack
Run A/B on real workloads: vendor tables are directional — blind-test against Kimi K3 or your current model on production Agent pipelines before full migration
Track official changelog for V4-Pro GA and Harness release: ignore unconfirmed August 10–20 rumors until DeepSeek's account or api-docs confirm dates; watch peak-hour 2x surcharge timing
If you plan to wire V4-Flash into Claude Code, OpenCode, or OpenClaw for long-session Agent workflows, running CLI tools on a laptop or unstable Linux VPS often means memory pressure, dropped sessions, and missing Xcode/Metal toolchains. For production environments that need stable SSH sessions, DerivedData caching, and iOS CI/CD automation — or local ds4 inference on a 96GB Mac — NodeMini Mac Mini cloud rental is usually the better fit: dedicated nodes with second-scale provisioning so Agent and build tasks run on a real Mac continuously. See Mac Mini rental rates and the help center. For local DeepSeek V4 deployment on rented Mac hardware, see our guide on ds4 inference on 96GB UMA Mac rental.
Sources: DeepSeek official API docs and changelog (api-docs.deepseek.com), DeepSeek-V4 technical report and Hugging Face model cards, Artificial Analysis (via Wantrich, Meyka), 21st Century Business Herald, Kuai Technology / ifeng Tech, V2EX community discussion, Moonshot AI, Zhipu/Z.ai, and Alibaba Cloud official announcements. Verify latest pricing and release status before publishing.
Yes. Both V4-Pro and V4-Flash, including the July 31 official V4-Flash-0731 build, ship as open weights under the MIT license on Hugging Face, and can be used, fine-tuned, and redistributed commercially without additional permission.
Based on figures reported by 21st Century Business Herald, official V4-Flash pricing runs roughly 36x cheaper than Claude Opus 4.8 on cache-miss input, about 179x cheaper on cache-hit input, and about 89x cheaper on output, per million tokens. These are vendor list prices, not an independent audit.
There's no confirmed date. DeepSeek's own changelog says only that the official V4-Pro release will follow as soon as possible. Reports of an August 10–20 general-availability window come from unnamed sources in Chinese media and haven't been confirmed by DeepSeek.
Partially. Widely-adopted third-party benchmarks like SWE-bench Verified carry more weight. But agent-specific scores (Terminal Bench 2.0, Toolathlon, etc.) were measured with DeepSeek's own unreleased Harness framework, and the company itself warns these numbers are highly sensitive to harness choice — wait for independent reproduction with other agent tools before treating them as general capability claims.
It's DeepSeek's first self-developed agent execution framework, positioned as an in-house alternative to Claude Code, for tasks like file editing, tool calls, and multi-step engineering work. It was named for the first time in the July 31, 2026 changelog and is not yet publicly available. For deployment environment questions, see help center or Mac Mini rental rates.