[SMM Analysis]Computing Power Providers Shift from Leasing to Tokens Amid Margin Squeeze and Soaring Demand

Published: Aug 18, 2026 09:36
Since Jul 2026, Chinese computing providers are shifting from GPU leasing to token billing—some call leasing "a side business." Leasing margins shrink under public price anchors (4090 8-GPU: CNY 7,200/mo) and broker markups; DeepSeek's Aug 17 API hike (3-5x) widens reseller spreads; daily token calls hit 140T in Mar 2026, up 1,000x in 2 years. SMM: margin recovery, not substitution—high-end GPU leasing stays tight; converts: small IDCs, resellers, hyperscalers.

Since July 2026, multiple computing power service providers contacted by SMM have been shifting their operational center from bare-metal leasing to Token-based billing services, with some bluntly stating that “computing power leasing is secondary.”On one side, leasing prices were flattened by publicly visible anchor points, and layers of intermediaries’ markups eroded shippers’ gross margins; on the other side, daily average Token call volume grew by more than 1,000x over two years, and official API price hikes widened the arbitrage space for transit/reselling—behind the migration in billing units from “selling card-hours” to “selling Tokens” lies the dual reality of leasing gross margins hitting bottom and a surge in inference demand.


I. Phenomenon: Multiple service providers say “leasing is secondary,” and Tokens have become the primary focus

According to SMM, since July, computing power service providers reached through multiple channels have proactively scaled back leasing matchmaking and shifted to Token businesses, with the shift intensifying month by month:

An IDC service provider in east China(three self-built data centers): explicitly stated that it is currentlyprimarily promoting Tokens, with computing power leasing as secondary; it has launched its own Token platform and disclosed a complete discount structure—40% of list price from midnight onward at night, different discounts for different models, and “the more popular the model, the lower the discount”; cloud producers’ Token discount ranges are about 46%–80% of list price.

A computing power service provider: upon receiving card-supply matchmaking information, it replied directly that its “center is on the Token model direction,” and it no longer undertakes leasing matchmaking.

An intelligent computing cloud service provider: its Token business adopts a transit agency model, purchasing at50%–60%of the original producer’s quoted prices and reselling; its current clients are mainly small-B and C-end users, and the business has just gotten started.

An operator: launched a Token private deployment model, providing key clients with secure dedicated lines directly connected to data centers; key-client discounts are 8.5–80% of list price, with price adjustments about once a month.

A leading cloud producer: channel feedback indicated that its center has already shifted to Tokens, and it is “not that interested” in the card-leasing business.

SMM’s direct impression from routine communications was that many service providers are no longer really doing leasing and are moving toward Tokens.The market has even seen partners targeting “Token revenue sharing”—bringing clients and model resources to seek idle computing power, with a cooperation starting point of 200 million TPM (Token calls per minute), and explicitly stating that “cooperation based on a monthly-rental model won’t work.”

Chart-1: Panorama of the Token discount structure: tiered pricing from 40% of list price at night to 8.5% discounts offered by operators


II. Supply side: Leasing gross margins were flattened by price anchoring, while Token pricing offers greater autonomy

The primary force driving the shift was the continued compression of profit margins on the leasing side.

Leasing Side—“Wholesale” Logic, with Monthly Revenue Capped. Taking a 4090 eight-GPU server as an example, the current market transaction anchor was about 7,200 yuan per unit·month. In primary brokerage, each resale added about 300-400 yuan per unit·month; after 2-3 flips, the quoted price could be pushed up to 8,500 yuan. However, the price anchor was public and transparent, and the profit from the markup stages went to the brokerage chain, leaving limited gross margin on the asset-owner side. Meanwhile, the market lacked standardized unified quotations; brokers accounted for a high share, and true asset owners had limited bargaining power in price negotiations.

Chart-2: 4090 Eight-GPU Monthly Rent—Under Price Anchoring, Divergence Between Broker Markups and Asset-Owner Gross Margin

Token Side—“Retail” Logic, with Revenue Tied to Usage. Token pricing mechanisms were flexible: 60% off starting at midnight, discounts fluctuating with model popularity, and monthly repricing in line with the market. For the same server, monthly rental revenue was capped, while usage-based Token billing offered greater per-GPU output elasticity during demand surges; off-peak valley-period discounts could also “shave peaks and fill valleys,” raising utilization at full load.

SMM Viewpoint: Leasing pricing, once supply became homogeneous, was anchored by the market; Token pricing units were “actual inference output,” with pricing power in the hands of the service providers themselves—this represented a shift in business logic from selling resources to selling output, and the gross-margin structure was reconstructed accordingly.


III. Price Catalyst: Official API Price Increase Implemented, Expanding the Transit Arbitrage Window

DeepSeek officially implemented a new round of API price adjustments starting August 17, with core models up 3-5x and some items reaching 10x; the Pro version output price rose from 6 yuan per million tokens to 27 yuan. From the channel side, SMM learned that some operators interpreted the substance as a return to the original price level before the prior roughly 2.5-discount sales promotions, rather than an entirely new price hike.

After the increase was implemented, the price spread between official listing prices and the 50-60% purchase price via transit channels widened, leaving a thicker arbitrage space for Token transit and agency models—this was also why the Token discount system had been repeatedly mentioned in channel communications since mid-August:service providers were seriously running the numbers on this.

Chart-3: DeepSeek Pro Output Price—Official Price Increase Expands the Transit Arbitrage Window

Public information showed that since 2026, API prices for high-capability models saw a phased rise:Zhipu’s Q1 API pricing increased 83%, while call volume still grew 400%; for Tencent Cloud’s Hunyuan series, prices for some models rose by more than 4x. Token pricing was shifting from a single price-cutting logic to tiered pricing: low-end Tokens were nearing commoditization and competing on price, while high-end Tokens were raising prices on the back of scarce capabilities; service providers with resale and integration capabilities benefited significantly.

SMM View: The original intent of the official price increase was to repair its own unit economics, but it objectively lifted the price-spread dividend for intermediary channels—every time the OEM adjusted prices, it created an opportunity for Token channel distributors to reprice.


IV. Demand Side: Daily Average Token Calls Rose by Over 1,000x in Two Years; Clients Wanted APIs, Not Machines

Industry data provided a footnote to this shift.Data from the National Data Administration showed that China’s daily average Token calls had increased from about 100 billion at the beginning of 2024 to more than 140 trillion in March 2026, a rise of over 1,000x in two years. At the WAIC 2026 conference this year, “Token factory” became the hottest key words, and the industry’s center extended from centralized training to scaled inference.

Chart-4: China’s Daily Average Token Calls: Leaping from the 100-Billion Level to the 100-Trillion Level

Supply-side follow-up was equally intensive:

  • all three major telecom operators had launched Token packages, selling Tokens in the form of data bundles; in May, Shanghai Telecom became the first local operator to release a Token tariff package, with 1 yuan corresponding to a 250,000-quota point allowance, and standard APIs enabling calls to more than 30 mainstream large models
  • After a listed computing-power enterprise upgraded its delivery model from computing-power leasing to billing based on actual Token usage, its H1 revenue was up 497% YoY, net profit was up 540% YoY, and it turned losses into profits
  • Some intelligent computing centers adopted an operating structure of “half full-lease wholesale, half Token retail”, with more than 2,000 users on the Token retail side
  • Outside China, there also emergedan inference infrastructure model of automated scheduling of idle GPUs plus Token revenue sharing, with monetizing idle computing power becoming a new entry point for the Token economy

SMM View: The structure on the demand side had changed—SMB and consumer clients did not care about GPUs and did not want bare metal; they only wanted an API key.Leasing service providers were stuck in the middle: they could not secure low-priced GPU supply upstream, while downstream clients no longer needed “an entire machine”. Tokens were the solution to this structural mismatch.


V. Boundary Observations: High-End GPU Leasing Remained Resilient, and the Shift Was Concentrated Among Three Types of Players

It should be noted that “switching from leasing to Tokens” was not an industry-wide phenomenon. The counterevidence SMM learned was equally clear: high-end card models such as the H200 remained in tight supply for leasing—including cases where suppliers breached contracts during the lease term and resold the capacity (with only one week’s notice in advance), a 128-card cluster leased as a whole at 116,000 yuan per card-month with no splitting into smaller orders, and a smart-computing center in Lingang where 2,000 domestic Muxi C500 cards were entirely subcontracted by two companies, who “still felt it wasn’t enough.”

Those that truly pivoted thoroughly fell into three categories: small and mid-sized IDCs and smart-computing cloud service providers without an advantage in card supply, channel players that started out as transit agents, and leading cloud providers that never relied on leased cards in the first place. The less tight a service provider’s cards were, the more thoroughly it pivoted; owners facing extreme scarcity still used long-term contracts to lock in scarce lease rents.


VI. Conclusions and Outlook: The Shift in Billing Units Is a Gross-Margin Repair, Not a Model Replacement

SMM believes: in this round, service providers’ shift toward Token is a resonance of push and pull forces—on the push side, lease-end price anchoring and the compression of card owners’ gross margins due to intermediary markups; on the pull side, exponential growth in inference demand, official price hikes widening the transit price spread, and higher revenue elasticity under usage-based billing; coupled with the fact that Token clients face high switching costs after connecting via API and have stickiness significantly stronger than lease clients who can move away at any time, the business logic of service providers shifting from “selling resources” to “selling output” is clear. For the same batch of servers, wholesale collects rent monthly while retail charges by Token—the contract has not changed; what has changed is the unit of measurement: one side says leasing gross margins have bottomed out, while the other says inference demand is surging.

Outlook:

Short term (3–6 months): the DeepSeek price-hike effect is expected to ferment; the dividend from the official-versus-transit price spread is expected to continue; expansion of Token transit and agency models is expected to accelerate; and small and mid-sized service providers are expected to continue launching Token platforms in rapid succession.

Medium term (6–12 months): tiered Token pricing is expected to deepen; commoditization of low-end Token is expected to intensify a channel reshuffle; if original manufacturers restart large-scale sales promotions and discounts, the room for transit arbitrage will be compressed, testing service providers’ client lock-in and value-added capabilities.

Long term (over 1 year): the core variable lies in high-end card supply—if NVIDIA card supply remains constrained and high-end card leasing stays a seller’s market, the dual-track structure of “collecting rent on tight-supply cards + selling Token on non-tight-supply cards” will run in parallel over the long term


Data sources

  1. National Data Administration — China’s daily average Token invocation volume data (about 100 billion times in early 2024 → surpassing 14 trillion times in March 2026)
  2. People’s Daily Online (2026-07-31), “The Industry Is Driving Computing-Power Services Into the ‘Token Era’” — WAIC 2026 “Token Factory,” and semi-annual report data on Token-based billing from listed computing-power enterprises
  3. Guojin Securities research report (2026-06) — Shanghai Telecom Token tariff packages, Token packages of the three major carriers, and an idle-GPU Token revenue-sharing model
  4. Public market information — Zhipu’s 2026 Q1 API pricing adjustment and call volume, and Tencent Cloud Hunyuan series price adjustments

Information obtained by SMM:

  1. Routine communication with an IDC service provider in east China — primarily promoting Token, with computing-power leasing as a supplement; Token discount structure (60% off at night, model-based floating, and 46–80% off for cloud providers)
  2. Routine communication with a smart-computing cloud service provider — Token transit model purchasing at 50–60% of the original manufacturer’s price, and a small B/C-end client mix
  3. Routine communication with a carrier — Token private-deployment model, key accounts at 85–80% of list price, and the DeepSeek price hike essentially being a restoration of the pre-2.5-discount original price
  4. Channel communication with a computing-power service provider — “the center is on the Token model direction”
  5. Channel feedback — a leading cloud provider’s center is on Token; the 4090 transaction anchor at 7,200 yuan per card-month, with intermediaries adding 300-400 yuan per lot; and the Lingang 2,000-card Muxi C500 fully subcontracted
 

Data Source Statement: Except for publicly available information, all other data are processed by SMM based on publicly available information, market communication, and relying on SMM's internal database model. They are for reference only and do not constitute decision-making recommendations.

For any inquiries or for more information, please contact: lemonzhao@smm.cn
For more information on how to access our research reports, please contact:service.en@smm.cn
Related News
[SMM Computing Power Express] 128 Units of 5090 Futures Delivered in Batches; Lessor Explores a Five-Year Long-Term Agreement
25 mins ago
[SMM Computing Power Express] 128 Units of 5090 Futures Delivered in Batches; Lessor Explores a Five-Year Long-Term Agreement
Read More
[SMM Computing Power Express] 128 Units of 5090 Futures Delivered in Batches; Lessor Explores a Five-Year Long-Term Agreement
[SMM Computing Power Express] 128 Units of 5090 Futures Delivered in Batches; Lessor Explores a Five-Year Long-Term Agreement
SMM learned that an operator in east China has been receiving deliveries of 128 units of RTX 5090 futures in batches. Since late June, SMM has been tracking the delivery status; to date, contract performance has been normal, with more than 40 units delivered. The expected price for a three-year contract lease is 13,000 yuan per unit per month. The lessor is also exploring a five-year long-term agreement model. SMM believes that the stable performance of 5090 leasing futures, together with the exploration of five-year long-term agreements, indicates that suppliers are promoting the evolution of computing power leasing toward quasi-commoditization and longer-term lock-in.
25 mins ago
[SMM Computing Power Express] 64 H100 Eight-GPU Servers in North China Available for Lease, Monthly Rent 75,000, Four-Year Closed-Mouth Term
2 hours ago
[SMM Computing Power Express] 64 H100 Eight-GPU Servers in North China Available for Lease, Monthly Rent 75,000, Four-Year Closed-Mouth Term
Read More
[SMM Computing Power Express] 64 H100 Eight-GPU Servers in North China Available for Lease, Monthly Rent 75,000, Four-Year Closed-Mouth Term
[SMM Computing Power Express] 64 H100 Eight-GPU Servers in North China Available for Lease, Monthly Rent 75,000, Four-Year Closed-Mouth Term
SMM learned that 64 H100 eight-GPU servers were released in North China, with monthly rent at 75,000 yuan per unit (four-year fixed term) and 76,000 yuan per unit (five-year fixed term). Each server was configured with 8×80G H100 GPUs, an Intel 8468 CPU, 2 TB DDR5 memory, and a 400G network card. The compute cards were interconnected at high speed via an IB switch. SMM believed that the market entry of H100 high-spec clusters under a four-year fixed term reflected high-end training demand increasingly concentrating on long-duration price locking.
2 hours ago
[SMM Hashrate Express] Subleasing of Hashrate Rentals Is Widespread; New Entrants Shift Toward Existing Resources
19 hours ago
[SMM Hashrate Express] Subleasing of Hashrate Rentals Is Widespread; New Entrants Shift Toward Existing Resources
Read More
[SMM Hashrate Express] Subleasing of Hashrate Rentals Is Widespread; New Entrants Shift Toward Existing Resources
[SMM Hashrate Express] Subleasing of Hashrate Rentals Is Widespread; New Entrants Shift Toward Existing Resources
SMM learned that the server subleasing model continued to prevail in the market. Procurement prices for NVIDIA GPU servers fluctuated at highs and lacked a price anchor, with the procurement price of 5090 fan cards approaching 40,000 yuan per unit. Many new entrants to the computing power leasing sector shifted toward integrating existing resources rather than purchasing additional hardware. SMM believed that the lack of an anchor for procurement prices was forcing the supply side to activate existing capacity through subleasing, and that the industry had entered a phase of competition over existing resources.
19 hours ago
[SMM Analysis]Computing Power Providers Shift from Leasing to Tokens Amid Margin Squeeze and Soaring Demand - Shanghai Metals Market (SMM)