[SMM Analysis] DeepSeek Sets a "Kill Line", Reshaping Global Large Model Pricing Power

Published: Aug 11, 2026 13:12
DeepSeek V4 Pro triggered a global price war among large language models with its ultra-low pricing. OpenAI slashed prices by 80%, Anthropic pulled back from markets outside China, Microsoft is evaluating integration, and Nvidia’s share of the Chinese GPU market plunged from 95% to 8%. Chinese models accounted for 60% of total call volume on OpenRouter, leading US models for 14 consecutive weeks.

I. Event Overview

On April 24, 2026, DeepSeek officially released the preview version of the V4 series large models, including the flagship DeepSeek-V4-Pro (1.6 trillion total parameters / 49 billion activated parameters) and the lightweight DeepSeek-V4-Flash (284 billion total parameters / 13 billion activated parameters). Both natively support a 1-million-token context window and are fully open-sourced under the MIT license. The V4 series entered formal availability on July 20, and the V4-Flash-0731 official version API launched for public testing on July 31.

On the same day, DeepSeek announced a permanent 75% price reduction for the V4-Pro API. In cache-hit scenarios, the input price dropped as low as 0.025 yuan per million tokens, with output at 6 yuan per million tokens—a new low price among global cutting-edge models.


II. Technological Breakthroughs: Not Just About Being Cheap

The V4 series introduced three core technologies:

  1. Hybrid Attention Architecture (CSA + HCA) : Combines compressed sparse attention with heavily compressed attention. Under a 1-million-token context, inference compute is only 27% of V3.2, and the KV cache uses only 10%.
  2. Manifold-constrained Hyperconnection (mHC) : Replaces traditional residual connections, ensuring training stability for networks with tens of thousands of layers.
  3. FP4/FP8 Mixed-Precision Training and Inference : First-ever industry implementation. Reduces VRAM bandwidth pressure by 50% and doubles inference efficiency.

On SWE-bench Verified (coding ability), V4-Pro scored 80.6%, leading all open-source models and nearly matching Claude Opus 4.6's 80.8%. It scored 87.5% on MMLU-Pro and 85.0% on AIME 2025 math reasoning. In its technical report, DeepSeek admitted being “slightly behind GPT-5.4 and Gemini-3.1-Pro by about 3–6 months,” but considering the price gap, this is almost irrelevant.


III. Price Slaughter: The “Execution Line” for Global Large Models

The pricing of V4-Pro directly shattered the global large model pricing system. Below is a comparison of output prices for major models in July 2026 (in US dollars per million tokens):

Chart-1: Comparison of API Output Prices for Major Models in July 2026 (Logarithmic Scale)

The ex-China developer community has reached a consensus: For the vast majority of C-end and small-to-medium B-end scenarios, incurring dozens of times the inference cost for marginal performance gains has no commercial feasibility. For any competitor priced higher than DeepSeek, only two paths remain: proactively significantly cut prices, or completely abandon the mass developer market.

Chart-2: Core Benchmark Comparison (Radar Chart)


IV. Impact on Major Models

4.1 OpenAI: From Arrogance to Urgent Price Cuts

OpenAI was hit the most directly. In 2025, its annual revenue was $13 billion, total spending $34 billion, and net loss $39 billion, with about $8 billion being operating losses.

On July 30, 2026, just three weeks after the release of GPT-5.6, OpenAI urgently announced an 80% price cut for the Luna version. CEO Sam Altman publicly acknowledged that cost is the "biggest challenge in the industry" and stated that "OpenAI is happy to cut prices by 75%," which was interpreted by observers as a passive response to DeepSeek.

On the OpenRouter platform, weekly token consumption by Chinese models has exceeded 60%, while that of US models has fallen to under 20% [9]. Data from the US cloud platform Vercel shows that the usage share of DeepSeek models in the enterprise market surged from 1% in April to 17% in May.

4.2 Anthropic: High-Price Strategy Under Pressure

Claude Fable 5 is priced at $50 per million tokens of output, 57 times that of V4-Pro. In June 2026, Anthropic began banning users outside China from using Claude, proactively shrinking its market—interpreted as a move forced by excessively high computing costs to focus on high-ARPU markets.

4.3 Google: Computing Power Shortage

Google Cloud's backlog orders have reached $460 billion, and Gemini has begun implementing usage quotas for major clients such as Meta.

4.4 Microsoft: Signals of Defection

On June 16, 2026, Microsoft officially began evaluating the integration of DeepSeek V4 into Copilot Cowork to replace OpenAI and Anthropic models. CEO Satya Nadella publicly stated: "Cheaper models and broader choices will enhance society's trust in AI; the public will not tolerate a few models and companies monopolizing all learning." Given Microsoft's 27% equity stake in OpenAI, this statement shook the industry.

4.5 Nvidia: CUDA Moat Cracking

On the day of DeepSeek V4's release, it achieved full adaptation across eight major Chinese GPU producers: Huawei Ascend, Cambricon, Hygon, Moore Threads, Muxi, Kunlun Core, Alibaba T-Head, and Tianjic. Bernstein predicts that Nvidia's share of the AI accelerator market in China will drop from 95% to around 8% in 2026, while Huawei Ascend's will rise to 50%–62%.

Chart-3: Key Events Timeline of the 2026 Large Model Price War


V. Capital Frenzy: 500 Billion Valuation

In June 2026: completed the first round of external financing exceeding 50 billion yuan, with a valuation exceeding $50 billion. Investors included Liang Wenfeng personally contributing 20 billion yuan, Tencent contributing 10 billion yuan, CATL contributing 5 billion yuan, JD.com contributing 3 billion yuan, and the National AI Industry Investment Fund.

August 2026: Launched a second round of 50 billion yuan financing, with a pre-money valuation of about 500 billion yuan (around $74 billion).

The total valuation of China's four major pure AI labs has reached about $159 billion, up from near zero 18 months ago.


VI. Market Data: Chinese Models Dominate Global Usage Rankings

Chart-4: Global Token Calling Share on OpenRouter Platform (July 27 – August 2)


VII. Reversals and Concerns

7.1 Computing Power Crunch, Prices About to Increase

On August 6, DeepSeek announced it was about to "significantly raise API prices." Its current computing power is equivalent to only about 20,000 H100 GPUs, far behind US competitors such as OpenAI.

7.2 Political Risks

The US government is highly vigilant about DeepSeek's models entering US enterprise infrastructure. Even when deployed in Microsoft's Azure US data centers, they still face security reviews.

7.3 Commercialization Paradox

The V4-Pro's price-to-sales (P/S) ratio exceeded 3,800 times, dozens of times that of Anthropic, but its annualized revenue is extremely low. Liang Wenfeng's pricing logic was that "recouping the server cost in 10 months is considered reasonable profit," and investors face the risk of valuation inversion.


VIII. SMM Commentary

SMM commented on this. First,the global large model pricing benchmark has shifted from a tacit understanding among the three Silicon Valley oligarchs to the "killing line" defined by DeepSeek. Meanwhile, the extreme efficiency route of MoE + sparse attention + mixed precision demonstrates that cutting-edge performance does not have to rely on stacking computing power. On the ecosystem side, domestic chips such as Huawei Ascend have gained large-scale practical opportunities, and the CUDA barrier is being shaken,The endgame of global large model competition may not be won by the strongest model, nor by the cheapest model, but by the optimal solution on the "capability/cost" cost-performance curve, which will be the ultimate winner.


References

  1. DeepSeek V4 Pro Arrives With 1M-Token Context, Reshaping Our Cost Calculus — —
  2. AI Wiki
  3. DeepSeek Becomes the Global AI Killing Line
  4. DeepSeek-V4-Pro: The Open-Source Model That's Rewriting the Rules of AI
  5. DeepSeek and Zhipu Break Through Computing Power Blockade: Zhipu Model Enters Top Three, DeepSeek Reduces Cost to 1/20
  6. How DeepSeek and Alibaba Are Shattering AI Price Barriers: A 70x Cost Differential Is Reshaping the Global Industry Landscape
  7. After the Official Release of DeepSeek-V4-Flash, Leading Model Producers Such as OpenAI Slashed Prices
  8. AI Inference Price War Explained
  9. DeepSeek's Inference Cost Drops to 1/36 of GPT-4o, How Will It Rewrite Industry Rules?
  10. AI Infrastructure Has Become the Kind of Business Munger Hates Most
  11. Microsoft Evaluates Using DeepSeek V4 Model to Cut Costs and Boost Efficiency, Potentially Replacing OpenAI and Anthropic
  12. AI Price War 2026: Why the Real Winner Isn't DeepSeek or OpenAI
  13. Google rations Gemini for Meta: when compute becomes the rarest resource
  14. Microsoft Eyes DeepSeek V4 for Copilot Cowork: What Azure Hosting Cannot Fix
  15. Domestic Chips Boom: What Nvidia Fears Most Has Arrived
  16. The Great Migration of AI Computing Power: 2026, Three Simultaneous Shifts
  17. DeepSeek Raises 7.4Billionat7.4Billionat50 Billion Valuation in First Outside Funding Round
  18. Reports: DeepSeek Has Restarted a Second Round of 50 Billion Yuan Funding
  19. Valuation as High as 500 Billion Yuan, DeepSeek Tops Global Large Model Call Rankings
  20. Low Prices Are Over! DeepSeek Set to Raise Prices Significantly — \
  21. What Are the Future Development Prospects of DeepSeek?

Data Source Statement: Except for publicly available information, all other data are processed by SMM based on publicly available information, market communication, and relying on SMM's internal database model. They are for reference only and do not constitute decision-making recommendations.

Images in this article contain AI-translated captions for reference only.

For any inquiries or for more information, please contact: lemonzhao@smm.cn
For more information on how to access our research reports, please contact:service.en@smm.cn
Related News
Brazil Data Center Grid Connection Applications Reach 38 GW, Attracting Over 206.7 Billion Investment
2 hours ago
Brazil Data Center Grid Connection Applications Reach 38 GW, Attracting Over 206.7 Billion Investment
Read More
Brazil Data Center Grid Connection Applications Reach 38 GW, Attracting Over 206.7 Billion Investment
Brazil Data Center Grid Connection Applications Reach 38 GW, Attracting Over 206.7 Billion Investment
According to official 2026 data from Brazil’s Ministry of Mines and Energy, the country’s national electricity system operator has received data center grid connection applications totaling 38 GW, corresponding to direct investment exceeding 206.7 billion yuan, which are planned to be implemented gradually over the coming years. Currently, there are 205 data centers in operation or under construction in Brazil. Leveraging its abundant renewable energy advantages and continuously updated policy incentives, Brazil is rapidly emerging as a core hub for global computing power capital.
2 hours ago
[SMM Computing Power Midday Review] The lowest price of the 4090 in the Yangtze River Delta rises, and bargaining power for inference computing power in east China shifts to the supply side.
5 hours ago
[SMM Computing Power Midday Review] The lowest price of the 4090 in the Yangtze River Delta rises, and bargaining power for inference computing power in east China shifts to the supply side.
Read More
[SMM Computing Power Midday Review] The lowest price of the 4090 in the Yangtze River Delta rises, and bargaining power for inference computing power in east China shifts to the supply side.
[SMM Computing Power Midday Review] The lowest price of the 4090 in the Yangtze River Delta rises, and bargaining power for inference computing power in east China shifts to the supply side.
In the Yangtze River Delta, the lowest monthly rental for a 4090 rose to 7,200 yuan (up 4.35%), and the lowest per-card hourly rate rose to 1.25 yuan (up 4.17%). As the mainstream consumer-grade AI inference card, the 4090 saw low-priced resources in the region continuously absorbed, pushing up lower-end prices and narrowing the quotation range. Bargaining power tilted toward the supply side, and sporadic supply reductions dominated.
5 hours ago
[SMM Computing Power Express] A certain intelligent computing cloud service provider is quoting monthly rental of 7,200 yuan/unit/month for 4090 eight-card servers, higher than the previous low quotation.
5 hours ago
[SMM Computing Power Express] A certain intelligent computing cloud service provider is quoting monthly rental of 7,200 yuan/unit/month for 4090 eight-card servers, higher than the previous low quotation.
Read More
[SMM Computing Power Express] A certain intelligent computing cloud service provider is quoting monthly rental of 7,200 yuan/unit/month for 4090 eight-card servers, higher than the previous low quotation.
[SMM Computing Power Express] A certain intelligent computing cloud service provider is quoting monthly rental of 7,200 yuan/unit/month for 4090 eight-card servers, higher than the previous low quotation.
SMM learned that a smart cloud computing service provider recently completed shipments of RTX 4090 8-card servers at a transaction price of 7,200 yuan per unit per month. Previously, the market had seen a relatively low quotation of 6,900 yuan per unit per month. The latest external price is about 300 yuan higher than that low point, indicating that although entry-level models are in ample supply, distributors have not continued to lower their external quotations. SMM believes that the stabilization of external quotations for the 4090 reflects that demand in the low and mid-end computing power segment has not yet experienced a cliff-like contraction.
5 hours ago
[SMM Analysis] DeepSeek Sets a "Kill Line", Reshaping Global Large Model Pricing Power - Shanghai Metals Market (SMM)