Why Did DeepSeek Raise API Prices by Up to 1,100%
Right After China's Open-Weight AI Blitz?

In a five-day window, three of China's top AI labs made moves that look contradictory on the surface. DeepSeek raised API prices by as much as 1,100% on certain tiers. Alibaba open-weighted a 2.4-trillion-parameter flagship it had never released before. Zhipu shipped GLM-5.3, boosting coding benchmarks by roughly 6x on the same base model — no retraining. This piece walks developers and buyers through the timeline, the rate card, the three strategies, the head-to-head, the disputed claims, and the FAQ. The shared signal: China's labs are shifting from competing on price alone to competing on pricing power.

01

Timeline: what happened, and when

The easy mistake is treating "1,100%" as the whole invoice. Start with the dates, then keep billing dimension, license type, and benchmark source separate.

DateEvent
Jul 16, 2026Moonshot AI open-weights Kimi K3 (2.8T parameters), drawing US security scrutiny
Aug 2–3, 2026Alibaba previews, then launches, Qwen3.8-Max as a hosted API
Aug 10, 2026Meta releases Muse Glimmer (30B, Apache 2.0), teases open weights for flagship Muse Spark 1.2
Aug 12, 2026Alibaba publishes Qwen3.8-2.4T-A95B open weights on Hugging Face / ModelScope; xAI ships Grok 4.6
Aug 13, 2026DeepSeek-V4-Pro goes GA and announces a price increase effective Aug 17; Google ships discounted Gemini 3.7 Flash
Aug 14, 2026Zhipu ships GLM-5.3, reusing GLM-5.2's 743B base
Aug 17, 2026, 00:00 Beijing timeDeepSeek's new pricing takes effect

Zoom out: on Jul 30, OpenAI cut prices on its cheapest tier (GPT-5.6 Luna, down 80%), then on Aug 6–7 made Luna the free default with unlimited text chats. While Chinese labs were raising prices and opening flagship weights, US labs were cutting prices and going free at the consumer layer — at the same time. That is two sides of the same pricing fight.

Six misreads to drop before you quote a headline

  1. 01

    Treat 1,100% as the whole bill. That figure is peak-hour cache-hit input, the tier that started nearest to free.

  2. 02

    Ignore Beijing peak windows. Peak hours are 9am–12pm and 2pm–6pm Beijing time. The announcement is asking you to shift load.

  3. 03

    Call any open checkpoint Apache 2.0. Qwen3.8-Max ships under a custom license. Muse Glimmer is a different bet.

  4. 04

    Repeat the geo-ban rumor. The published license has no US / EU / UK / Korea territorial clause.

  5. 05

    Read GLM-5.3 as a new foundation model. Same 743B base as 5.2. The jump is post-training RL scale.

  6. 06

    Assume Chinese model still means cheapest model. Off-peak DeepSeek still undercuts Claude Opus 5. Luna and Qwen international pricing now undercut DeepSeek.

02

The numbers: what actually changed

Prices below come from DeepSeek's official announcement, cross-checked against Wall Street CN, IT Home, and V2EX. GLM scores are Zhipu's own reported numbers. No independent third-party re-run has been published yet.

DeepSeek's price hike, tier by tier

Effective Aug 17, 00:00 Beijing time. Peak hours: 9am–12pm and 2pm–6pm Beijing time. Unit: per 1M tokens.

Billing itemOld priceNew off-peakNew peakPeak increase
V4-Flash cache hit (input)¥0.02¥0.05¥0.10~400%
V4-Flash cache miss (input)¥1.0¥1.5¥3.0200%
V4-Flash output¥2.0¥4.5¥9.0350%
V4-Pro cache hit (input)¥0.025¥0.15¥0.30~1,100%
V4-Pro cache miss (input)¥3.0¥4.5¥9.0200%
V4-Pro output¥6.0¥13.5¥27.0350%
info

Context: The headline 1,100% figure applies to peak-hour cache-hit input. Output pricing, which dominates most real bills, rose 350%. Independent cost modeling found a realistic heavy-usage workload (roughly 84M tokens/month, mostly off-peak, half cache hits) sees a bill increase closer to 1.8x.

Qwen3.8-2.4T-A95B (Qwen3.8-Max open weights)

SpecDetail
Parameters2.4T total, 95B active per token (MoE, 512 experts, 10 routed + 1 shared)
Context window262,144 tokens native (open checkpoint), extendable to ~1.01M; hosted Max version defaults to 1M
Release cadencePreview Aug 2 → API live Aug 3 → open weights Aug 12
API pricing (international)$2/M input, $6/M output
LicenseNot Apache 2.0 — a custom Qwen3.8-Max License
Why it mattersFirst time Alibaba has open-weighted a Max-tier flagship; Qwen3.5 / 3.6 / 3.7 Max stayed API-only

GLM-5.3 vs GLM-5.2: same base, post-training only

BenchmarkGLM-5.2GLM-5.3Change
Terminal-Bench 3.04.6%28.3%+23.7 pts
DeepSWE v1.146.2%66.9%+20.7 pts
Agents' Last Exam (CLI)23.8%28.5%+4.7 pts
CyberGym77.2%84.5%+7.3 pts
AutomationBench26.2%48.2%+22.0 pts
warning

Caveat: These are Zhipu's own reported numbers. GLM-5.3 still trails GPT-5.6 Sol (34.6%) and Claude Fable 5 (33.7%) on Terminal-Bench 3.0. It is a top open-weight result, not an outright frontier win.

03

Breaking down the three strategies

Three labs. Three plays. One question underneath: when usage scales faster than GPUs and pretraining returns flatten, who still holds pricing power.

DeepSeek: time-of-day pricing is a capacity problem

The easiest misread is "China's cheapest model finally caved to margin pressure." The structure reads more like the opposite: a company making compute constraints visible in the price sheet. Flat, always-cheap pricing worked as a customer-acquisition tool as long as GPU capacity kept pace. Once usage grew exponentially and capacity did not, something had to become explicit. "Encouraging more flexible workload scheduling" is corporate-speak for "peak-hour compute is now scarce, please shift your load yourself."

One detail international coverage mostly missed: at peak hours, DeepSeek's official API price is now higher than several third-party resellers (GMI Cloud, Novita, and others currently list V4 Pro below DeepSeek's new peak rate). The assumption that the official API is always the cheapest way to run DeepSeek has been broken for the first time.

Alibaba: open weights buy mindshare; a custom license protects the ceiling

Alibaba did two things at once. It published the full 2.4T-parameter checkpoint for free download, and it attached a custom license — not the permissive Apache 2.0 used for smaller Qwen models. Any Model-as-a-Service or AI Work Assistant business earning over $50 million in any 12-month period must negotiate a separate commercial license. Products with 100M+ monthly active users or $20M+ in monthly revenue must prominently display the model name.

Give away the weights to win developer mindshare, especially internationally. Keep pricing leverage over the handful of companies that can build a competing inference business on top. That is a materially different bet than Meta's Muse Glimmer, which ships under unrestricted Apache 2.0.

info

Rumor to kill: Claims that Alibaba's license bans downloads from the US, EU, UK, and South Korea are false. The published license text contains no geographic or territorial clause of any kind. Check the LICENSE file, not the announcement thread.

GLM-5.3: no new base model, just a bigger post-training bet

The most interesting fact is the method: same 743B-parameter base as GLM-5.2, no retraining, and a roughly 6x jump on Terminal-Bench 3.0 (4.6% → 28.3%) from scaling reinforcement-learning environments in post-training. As pretraining scaling laws show diminishing returns, post-training RL scale is becoming an independent performance lever with a much lower cost floor than retraining a new foundation model. Mid-tier labs without OpenAI-scale compute budgets can still close the gap on agentic and coding benchmarks.

China's AI labs are shifting from competing on price alone to competing on pricing power itself.

04

Head-to-head: is DeepSeek still the cheapest frontier-class model?

RMB-to-USD conversion at about ¥7.15/$1, approximate. Official announcements win if they diverge.

ModelInput (per 1M tokens)Output (per 1M tokens)Open weights?
DeepSeek V4-Pro (peak)¥9.0 (~$1.26)¥27.0 (~$3.78)No
DeepSeek V4-Pro (off-peak)¥4.5 (~$0.63)¥13.5 (~$1.89)No
Qwen3.8-Max (international API)$2.00$6.00Yes (custom license)
OpenAI GPT-5.6 Luna$0.20$1.20No
Claude Opus 5 (implied, per Alibaba's comparison ratio)~$5.00~$25.00No

The short answer: no. DeepSeek V4-Pro's off-peak rate is still well below Claude Opus 5, but it is no longer the outright cheapest option. Qwen3.8-Max international pricing and OpenAI's Luna now undercut DeepSeek's off-peak rate. "Chinese model = cheapest model" was true for most of 2025 and early 2026. It is not a safe assumption anymore.

Six steps after the hike: invoice, routing, license

  1. 01

    Split the invoice into three columns. Cache-hit input, cache-miss input, and output. Do not budget from a single 1,100% headline.

  2. 02

    Map jobs to Beijing peak hours. 9am–12pm and 2pm–6pm Beijing time bill at peak. Batch what you can into the other 17 hours.

  3. 03

    Measure cache-hit rate. A modeled heavy workload (~84M tokens/month, mostly off-peak, half hits) lands near 1.8x, not 12x.

  4. 04

    Compare official peak to resellers. GMI Cloud and Novita currently list V4 Pro below DeepSeek's new peak. Official is no longer automatically cheapest.

  5. 05

    Read the LICENSE file before commercial use. The Qwen triggers are $50M MaaS / AI Work Assistant revenue and 100M MAU or $20M monthly revenue for attribution. No geo ban.

  6. 06

    Accept GLM-5.3 as a post-training lever. Official Terminal-Bench 3.0: 28.3% vs 4.6%. Still behind Sol 34.6% and Fable 5 at 33.7%. Validate on that basis.

text
# DeepSeek V4-Pro output (per 1M tokens, official card)
Old flat rate:   ¥6.0
New off-peak:    ¥13.5   (+125% vs old; half of peak)
New peak:        ¥27.0   (+350% vs old)

# Three questions before you quote a headline
1) Cache-hit input, cache-miss input, or output?
2) Peak or off-peak?
3) What is my real cache-hit rate?
05

What is disputed, why it matters, and three numbers you can cite

What's disputed or unverified

  • The 1,100% headline is technically accurate but misleading without context. It applies only to peak-hour cache-hit input. Output pricing rose 350%. Different outlets have quoted different tiers as if they were the whole story.
  • Claims that Qwen3.8-Max runs on Alibaba's in-house Zhenwu M890 chips (and "Pangu AL128" supernodes), reported by several Chinese financial outlets, have not been independently confirmed by Alibaba's own technical documentation or third-party benchmarks. Treat this as vendor-adjacent, unverified reporting.
  • GLM-5.3's reported discovery of a "serious vulnerability" in Cursor comes from VentureBeat and Zhipu's own disclosure. Specific technical details have not been made public. Read it as a vendor-sourced, not independently audited, security finding.
  • Reports that China's Ministry of Commerce may be preparing retaliatory export controls on AI / semiconductor technology are speculative and sourced to unconfirmed media reports, not an official announcement.

Two price wars running in parallel

Over roughly the past month, China's top labs have shipped major releases at a pace domestic financial media has started calling "three model updates a week" (一周三更) — DeepSeek, Alibaba, and Zhipu, plus Moonshot's Kimi K3 (open-weighted Jul 16, 2.8T parameters) and MiniMax H3. Chinese coverage broadly frames this as Chinese open-weight releases forcing a global repricing of the AI industry.

US labs are running the opposite play at the consumer layer: OpenAI cut prices 80% on its cheapest tier (Jul 30) then made that model free and unlimited a week later (Aug 6–7); Google shipped a coding-focused model at half the price of its three-week-old predecessor (Aug 13). Chinese labs open-weight flagships and introduce tiered, higher pricing on the compute-constrained top end. US labs race toward free and cheap at the consumer end.

There is also a geopolitical layer worth naming carefully. Moonshot's Kimi K3 open-weighting in July already drew US security scrutiny. Alibaba choosing this window to open-weight a 2.4T flagship has been read by some analysts as a move to lock in international mindshare and a "technological parity" narrative before any potential regulatory tightening. That is an informed interpretation, not a confirmed fact — and it is easy to miss if you only read English-language product posts.

Citeable figures (as of publication)

  • DeepSeek V4-Pro peak output: ¥27.0 per 1M tokens, up 350% from ¥6.0. Cache-hit input: ¥0.025 → ¥0.30, about 1,100%.
  • Qwen3.8-2.4T-A95B: 2.4T total / 95B active, MoE 512 experts (10 routed + 1 shared), 262,144-token native open checkpoint.
  • GLM-5.3 Terminal-Bench 3.0: official 28.3% (5.2 was 4.6%), still below GPT-5.6 Sol 34.6% and Claude Fable 5 at 33.7%.

Sources: DeepSeek's official pricing announcement, cross-checked against Wall Street CN, IT Home, AIGC.cn, and V2EX; Alibaba's official Qwen repositories (Hugging Face / ModelScope) and South China Morning Post on license terms; Zhipu (Z.ai) GLM-5.3 technical page, plus VentureBeat and StableLearn; Meta AI Research's official blog; Chinese financial outlets (Yicai, Sohu Finance) on release pacing. Verify the latest official pricing and license terms before republishing. Details flagged above as unverified have not been independently confirmed.

Peak/off-peak API pricing makes "always-on flagship inference" more expensive and harder to forecast. Self-hosting a 2.4T MoE is a datacenter problem, not a laptop problem. Resellers can shave the peak rate, but parking production agents, Xcode builds, and long sessions on a shared inference queue still leaves you with weak isolation. For a more stable production environment that fits iOS CI/CD and AI-agent automation, NodeMini's Mac Mini cloud rental is usually the better fit — dedicated Apple Silicon, node-level billing, no fight with Beijing peak hours for the same GPU invoice.

FAQ

Frequently asked questions

Its off-peak rate is still cheaper than Claude Opus 5, but it is no longer the single cheapest option overall — OpenAI's GPT-5.6 Luna ($0.20/$1.20 per million tokens) and Alibaba's international Qwen3.8-Max pricing ($2/$6) now undercut DeepSeek's new off-peak rates on at least one dimension. If you want coding agents off the shared peak-hour queue, compare Mac Mini rental rates.

Yes, for most use cases — personal projects and internal enterprise use are unaffected. The catch applies only if you are running a Model-as-a-Service or AI Work Assistant business that has earned over $50 million in any consecutive 12-month period; that tier requires a separate commercial license from Alibaba.

No. That claim circulated online but is false — the published license contains no geographic restriction of any kind. The restrictions are revenue-based (tied to how much money your service makes), not tied to where you or your users are located.

Nothing at the base-model level — both use the same 743-billion-parameter foundation model. The performance gains (roughly 6x on Terminal-Bench 3.0) come entirely from scaling up reinforcement learning during post-training, with no retraining of the base model.

Not yet. Muse Glimmer is a 30B distilled model, not Meta's real flagship. CEO Mark Zuckerberg has said open weights for the larger, closed Muse Spark 1.2 are coming "soon." As of this writing, that release has not happened; treat it as a stated intention, not a confirmed fact. Isolation and access notes are in the help center.