Sino-US Token Economics: Profit Sources, Premium Flows, and Realization Sequence

Wallstreetcn
2026.08.09 09:59

Profits across the large language model (LLM) industry chain are being realized sequentially from upstream to downstream. The upstream computing infrastructure layer is the first to capture scarcity dividends, while the midstream model layer is mired in price deflation due to open-source parity. Regarding premium flows, a significant divergence exists between China and the US: In the US, incremental AI value accumulates within existing high-priced software subscription systems; in China, due to differences in willingness to pay, low-priced token dividends spill over directly to the downstream application layer. In the future, whether downstream vendors can retain profits will depend on the "switching costs" they build within their business scenarios

The mismatch between the surge in LLM computing power and sluggish monetization is reshaping the profit distribution landscape of the global AI industry chain. Over the past two years, daily token call volumes in the Chinese market have skyrocketed by more than a thousandfold, yet the annual revenue for public cloud Model-as-a-Service (MaaS) in 2025 remained at merely the RMB 3 billion level. Massive consumption has not yet translated into commensurate book revenue, and China and the US have taken distinctly different paths regarding computing bottlenecks and commercialization strategies.

In an analysis of the token economy industry chain, Song Xinzhu, an analyst at Northeast Securities, proposed that AI profit accumulation consists of four mechanisms: scarcity premiums, generational premiums, integrated internal settlement gains, and switching cost premiums.

Currently, profits are entering financial statements sequentially from upstream to downstream: the upstream computing infrastructure layer realizes scarcity dividends first; the midstream model layer is mired in deflation triggered by the commoditization of same-generation capabilities; and the downstream application layer captures the benefits of falling computing costs, building long-term moats through "switching costs" accumulated over time.

Regarding the ultimate direction of premium flows, constrained by differences in market payment endowments between the two countries, incremental AI value in the US is accumulating within high-priced software subscription systems, while in the Chinese market, low-priced token dividends spill over directly to the application layer, awaiting value reassessment after a comprehensive shift in pricing models.

Computing Investment Approaches Cash Flow Boundaries; Thousand-fold Call Volume Yields Only RMB 3 Billion Market

The token economy remains in a heavy-asset construction phase. On the demand side, daily call volumes in China surged from approximately 100 billion in early 2024 to 100 trillion by the end of 2025. However, most token consumption occurs within the proprietary scenarios of major tech firms, failing to form external transactions; for those parts involving external transactions, transaction prices have been extremely compressed. Furthermore, as application-layer charging has not fully shifted to token-based pricing, a thousand-fold increase in usage has resulted in only a RMB 3.07 billion public cloud MaaS market size.

Corresponding to the meager API revenue is an extremely heavy load on the computing investment side. Capital expenditure intensity is approaching the coverage boundary of operating cash flow. By the second quarter of 2026, the ratio of trailing twelve months (TTM) capital expenditures to operating cash flow for the four major US cloud providers rose to between 0.63 and 1.05. Alphabet recorded negative free cash flow for a single quarter for the first time, while Meta's free cash flow plummeted 91% year-on-year. Funding sources for the construction phase have spilled over from operating cash flow to capital markets.

Investment pacing in the Chinese market shows significant divergence. Alibaba's capital expenditure intensity ranks first among Chinese concept stocks, while Baidu is the only company among the top eight computing buyers experiencing both a decline in revenue and an increase in capital expenditure.

Upstream Captures Scarcity Dividends; Sino-US Computing Bottlenecks Diverge

The upstream sector is currently the only link steadily securing profits in its financial statements. The "scarcity premium" based on supply gaps directly created NVIDIA's USD 193.7 billion in data center revenue for fiscal year 2026. Facing the same hunger for computing power, China and the US have formed distinctly different clearance methods and industry bottlenecks under the same regulatory constraints.

The bottleneck in the US industry chain lies in power grid access. In the queue waiting for access approval at ERCOT (Electric Reliability Council of Texas), about 90% of the over 1,800 projects are data centers, corresponding to a cumulative electricity demand of approximately 474 GW. Extended approval and power access cycles have driven North American data center vacancy rates to historic lows. Scarcity in the US is ultimately cleared through price, with the benefits of price hikes accruing to head firms like NVIDIA.

The bottleneck in China's industry chain points directly to computing chips. Under export controls, the Chinese market clears according to controlled allocations, with institutional pull directing towards domestic substitution. In 2025, local vendors accounted for more than 40% of the AI accelerator card market. New computing capacity is concentrating in hub nodes of the "East Data, West Computing" project, with construction entities comprising public sectors, operators, and private capital, forming a public-sector-led computing system.

Open Weights Break Generational Barriers; Midstream Models Become Standardized Capacity

Tokens of the same capability tier possess strong substitutability, and open weights (open source) have become the absolute main force in leveling price disparities. The cost for buyers to switch suppliers is extremely low, so competition falls directly on listed prices. Calculations show that the call price for achieving GPT-4 equivalent capabilities drops to about one-fortieth annually.

Prices for same-generation capabilities are rapidly converging globally. At the AA Intelligence Index level of approximately 51 points, the blended prices of four major models from China and the US (GPT-5.6 Luna, GLM-5.2, MuseSpark 1.1, Gemini 3.6 Flash) all fall within the very narrow range of RMB 14 to 22 per million tokens. The lowest price in this tier does not come from a Chinese vendor, but from Meta, which entered via API. Once capability tiers are leveled by open source, tokens become commoditized, with prices changing only with volume and cost.

Consequently, the midstream is squeezed from both ends. On the selling side, actual transaction prices are often about an order of magnitude lower than listed prices; on the cost side, depreciation and amortization account for more than 70% of unit inference costs, while electricity costs make up less than 10%. The core space for cost reduction lies not in electricity prices, but in depreciation periods and computing utilization rates.

In China, token commoditization is being pushed to the infrastructure level. Through direct subsidies to buyers via computing vouchers and the establishment of unified measurement and comparison platforms, intermediate markup spaces are being squeezed out. The circulation link is destined for increased volume but thin margins; only the two ends can continue to make money: the production side earns cost differentials through extreme utilization rates, and the scenario side retains profits through customer switching costs.

Divergence in Profit Destination: Accumulation in Subscriptions in the US, Spillover to Application Layer in China

Deflating token prices release the dividends of model upgrades to the downstream, with China and the US adopting different ways to capture these dividends due to differences in willingness-to-pay endowments.

Incremental AI value in the US is absorbed by existing subscription systems. US users are accustomed to paying high prices for software subscriptions, and frontier model firms compete for the same ecological niche as incumbent software giants. Microsoft bundles AI features into high-priced subscriptions, converting one-time capability advantages into continuous customer payments.

The Chinese market is constrained by a lower willingness to pay for software subscriptions. For the same basic office software, pricing in the Chinese market is typically only one-fifth or even lower than in the US. Therefore, Chinese midstream firms generally price based on computing costs, and the value of low-priced tokens spills over directly to the application layer.

For Chinese application-layer companies, the shift in pricing models determines where profits go. If subscription pricing is maintained, token price reduction dividends remain on the cost side, reflected in improved gross margins and operating leverage; if the shift is towards token consumption or business outcome-based pricing, dividends enter the revenue side directly.

However, a change in pricing models does not equate to profits landing in the pocket. Whether downstream players can retain profits depends on the "switching costs" established within scenarios, which is the only weapon to resist buyers pressing prices to reclaim dividends.

The attributability of output value, the exclusivity of scenario assets, customer structure, and task consumption intensity determine the quality of the scenario. Moats within scenarios are maintained by three types of carriers: first, access qualifications and procurement systems granted by external rules (such as government affairs and regulated industries); second, data integration and standard accumulation formed over time (such as financial data governance and intelligent operations/maintenance); and third, reset costs brought about by system-level integration.

As the old seat-based billing system collapses, pricing models based on tokens and outcomes will significantly shorten the verification cycle for customer stickiness. Customer retention, which previously required waiting for annual renewals to determine outcomes, can now be clearly observed through quarterly token usage. Ultimately, those who can accumulate profits are the downstream players who bind customers to their own business flows with extremely high replacement costs.