At WAIC in Shanghai this month, across 1,100 exhibitors and more than 140 forums, the phrase that kept coming up was not AGI. It was Token工厂, token factory.
There is nothing subtle about the strategy behind it. Take a unit of AI output and standardize it. Industrialize the production. Drive the cost down until nobody else wants to compete, then export the surplus. It is the solar and battery playbook, applied to inference.
I run a cross-border business out of Shanghai, and before that spent seven years running a commerce agency here on a P&L that lived or died on Chinese platform economics. That makes me skeptical in two directions: of Western coverage that waves off Chinese infrastructure, and of Chinese coverage that treats a pilot as a finished export industry.
To be clear up front, none of this argues Chinese models are weak. Stanford’s AI Index put the top US model 2.7 percent ahead of the top Chinese model in March. The capability question is close to settled. What follows is about infrastructure and economics.
Five assumptions hold the strategy up. One stands. Four are shaky. If you only want the practical part, skip to the last section.
The assumption ledger — open a row to see what the available data does to it (1/5 stands)
| Assumption | Meter | Verdict | Detail |
|---|---|---|---|
| The export loop works | 22% | One pilot | Shantou closed the loop in April: data in, inference domestic, answer out by API. No volume and no revenue have been published, and the city was picked for a cable landing station most of the country does not have. |
| The price gap is real | 92% | Stands | Open-weight Chinese models run 60 to 90 percent below leading Western pricing, and developers have already routed accordingly. This is the strongest leg the strategy has. |
| Domestic silicon can carry it | 30% | Not yet | Huawei’s 910C delivers roughly 60 percent of H100 inference performance, and most of DeepSeek’s tokens are still inferenced on Western hardware. |
| The factories are running | 28% | Mostly idle | Effective usage of 36.8 percent in 2025, rack utilization of 20 to 30 percent at many centers, against leading platforms running at 90 to 95 percent. A coordination problem, not a supply glut. |
| Volume converts to revenue | 18% | Inverted | Anthropic holds around 12 percent of token volume on OpenRouter and captures close to 46 percent of the revenue. China owns the commodity lane; the profit sits in the other one. |
Legend: Survives the data · Running ahead of the evidence.
36Kr’s WAIC recap called it the first year the token economy became a core conference topic.
— 36氪, July 2026
Make inference a metered commodity
In March, China’s National Data Administration fixed the official Chinese term for token as 词元, cí yuán, and called it a settlement unit linking technical supply to commercial demand. Beijing now tracks daily token volume and reports it at policy events the way it reports steel output.
The headline number is a thousandfold rise in two years. About 100 billion daily calls in early 2024, more than 140 trillion by March 2026.
- 100B daily calls, early 2024
- 140T daily, March 2026 — a 1,000× rise
- State-reported. Unaudited. Counts consumption, not value.
Treat that with some care. It is state-reported, nobody audits it, and it counts consumption instead of value. That is not a China problem specifically. No government or vendor anywhere publishes an audited token number. An agent stuck in a bad loop overnight burns an absurd quantity and produces nothing.
By end-2025 China had built more than 100,000 high-quality datasets totaling over 890 petabytes, the supply side of the same policy push.
— 新华社, March 24, 2026
The mechanism is real, but it is one pilot
In April, Shantou in Guangdong closed a loop under a policy called 来数加工, inbound data processing. Overseas data enters a digital bonded zone. Domestic compute runs the inference. The answer goes back out by API. Power, compute and revenue all stay in China.
- Overseas data in — enters a digital bonded zone
- Domestic compute — runs the inference
- Answer out by API — only the output crosses back
Power stays. Compute stays. Revenue stays.
Chinese state media summed it up with a line that traveled widely: the salt can travel the world, but the salt fields must stay home. Good line. I have caught myself repeating it. It is also promotional framing from the same system that publishes the consumption figures.
What has not been published is volume. One city, one pilot, no disclosed revenue. Shantou was picked for reasons that predate AI entirely. It holds one of three mainland submarine cable landing stations, and latency to Singapore runs about 32.7 milliseconds. Most of the country has neither. That part tends to get left out.
On the record:
- 1 of 3 mainland submarine cable landing stations
- 32.7 ms latency to Singapore
Not published:
- n/a export volume
- n/a disclosed revenue
Chinese coverage frames token exports as an impossible triangle of latency, data compliance and very cheap green power.
— 澎湃新闻 and 同花顺财经, 2026
It is real, and developers have already voted
Nobody argues with this one. It is the strongest leg the strategy has.
DeepSeek V4 Flash costs $0.14 per million input tokens. GPT-5.5 costs $5.00. Open-weight Chinese models run 60 to 90 percent below leading Western pricing. On OpenRouter, which routes calls to whichever provider wins on price, Chinese models went from under 10 percent of usage in early 2025 to a record 58 percent of tokens processed by US firms in July.
| Model | Input, per 1M tokens |
|---|---|
| DeepSeek V4 Flash | $0.14 |
| GPT-5.5 | $5.00 |
Open weights run 60 to 90 percent below leading Western pricing.
Chinese model share on OpenRouter (of tokens processed by US firms): under 10% in early 2025, rising to 58% in July 2026, a record.
In agent workloads the gap compounds, because one task can fan out across dozens of parallel sub-agent calls.
OpenRouter weekly volume grew from roughly 5 trillion tokens in April 2025 to over 20 trillion a year later.
— CNBC, July 2026
Where it breaks first: the compute underneath
This is the gap most coverage understates, and it is the one that bothers me. You cannot industrialize output on equipment you cannot get enough of.
Huawei’s 910C delivers roughly 60 percent of H100 inference performance, by DeepSeek’s own published assessment. The Council on Foreign Relations puts the best US chips at about five times more powerful than Huawei’s best today, widening toward 2027. Then there is the SemiAnalysis detail, which matters more than any chip spec: most of DeepSeek’s tokens are still inferenced on Western hardware. The flagship of Chinese open-weight AI is not running the Chinese stack at scale.
| Comparison | Value | Note |
|---|---|---|
| Huawei 910C vs H100 | 60% | of H100 inference performance |
| Best US chip vs best Huawei | 5× | more powerful today, widening toward 2027 |
| CloudMatrix 384 vs NVL72 | 4.1× | the power, for about 1.7× the compute |
The flagship of Chinese open-weight AI is not running the Chinese stack at scale.
The energy argument, which usually gets made next, has the same problem. China generates more than twice the electricity of the United States and added 543 gigawatts in 2024 alone. But a four-fold power penalty on domestic accelerators eats much of a two-fold price advantage. Cheap power ends up subsidizing inefficient silicon rather than producing cheaper tokens.
The advantage: 2× the electricity of the United States, plus 543 gigawatts added in 2024 alone — eaten by the penalty: 4× the power draw on domestic accelerators, against a two-fold price advantage.
Chinese firms are not sitting still. Zhipu trains and serves on domestic accelerators. Alibaba is building data centers across South Korea, Malaysia, Thailand and Mexico, renting foreign capacity to sidestep the silicon problem. Both are rational moves, and neither closes the gap this year.
Huawei’s CloudMatrix 384 draws roughly 4.1 times the power of Nvidia’s NVL72 for about 1.7 times the compute.
— SemiconductorX, 2026
A lot of the capacity is sitting idle
The factory framing implies the plants are running. Chinese domestic reporting suggests many are not.
Tencent Cloud puts average GPU utilization at China’s intelligent computing centers below 30 percent. CAICT data cited by Huxiu shows an effective usage rate of 36.8 percent in 2025, meaning roughly 70 yuan of every 100 invested is idling. Securities Times reported rack utilization of 20 to 30 percent at many centers.
- Below 30% — average GPU utilization at intelligent computing centers (Tencent Cloud)
- 36.8% — effective usage rate in 2025 (CAICT, via Huxiu)
- 20 to 30% — rack utilization at many centers (Securities Times)
Calling this oversupply misses it. The market has split, and I have watched the same pattern in Chinese logistics: healthy national totals sitting on stranded regional capacity. Leading platforms run at 90 to 95 percent with order books into 2028. The high-quality compute is genuinely short, while the low-quality stuff cannot find a customer at any price. That is a coordination problem, and those take years to clear.
- High-quality compute — leading platforms run at 90 to 95 percent with order books into 2028. Genuinely short. (order books into 2028)
- Stranded capacity — the low-quality stuff cannot find a customer at any price. A coordination problem, and those take years to clear. (no customer at any price)
A thousand-card computing center costs over 30 million yuan a year to operate, which is why low utilization turns into losses fast.
— 虎嗅 and 观察者网, July 2026
Volume is not revenue, and that is the gap that matters
If you read one number here, make it this one. Anthropic holds around 12 percent of token volume on OpenRouter and captures close to 46 percent of the revenue.
| Metric | Anthropic share, on OpenRouter |
|---|---|
| Share of token volume | 12% |
| Share of revenue | 46% |
The lanes are priced differently.
The market is splitting into a commodity lane and a premium lane. China owns the commodity lane. The profit sits in the other one.
The economics are moving the wrong way for everyone. Huxiu documents Alibaba Cloud and Baidu AI Cloud raising prices by up to 34 percent this year. Agents consume 100 to 1,000 times what a chatbot does. Unit prices keep falling while invoices keep growing, and nobody has an accepted way to link tokens consumed to work completed.
- Up to 34% — price rises at Alibaba Cloud and Baidu AI Cloud this year
- 100 to 1,000× — what an agent consumes against a chatbot
- 85% — of enterprise AI budgets now going to inference
Inference now accounts for roughly 85 percent of enterprise AI budgets, against near-zero marginal cost in traditional software.
— 腾讯新闻, June 2026
The one no price cut touches
Under China’s 2017 National Intelligence Law, companies must cooperate with state intelligence work. That is a jurisdictional fact. It says nothing about how any particular company behaves, and it is still enough to keep regulated finance, healthcare and government workloads with Western providers.
Washington is escalating separately. The Treasury Secretary raised the prospect of sanctions in July over alleged distillation of US models, days ahead of a September bilateral AI dialogue. If part of the Chinese cost advantage comes from not paying frontier training costs, it is less durable than it looks. Nobody has proven that case. It is not trivial either.
The practical response The practical response is unglamorous. If you route production traffic through Chinese-origin endpoints, keep a tested fallback configured. Model availability became a policy variable this year, and it cuts both ways.
Large enterprises continue buying mostly from Anthropic, OpenAI, Azure and Google Cloud.
— Baiguan, July 2026
What would change my mind
I am not predicting failure. The export story is running ahead of the evidence, which is different. Three signals would move me: Shantou publishing real export volume and revenue, Chinese computing center utilization crossing 50 percent, and a regulated Western enterprise putting a production workload on a China-hosted endpoint.
- Shantou publishes real export volume and revenue.
- Chinese computing center utilization crosses 50 percent.
- A regulated Western enterprise puts a production workload on a China-hosted endpoint.
Any of those would tell you more than another month of routing charts. If all three move, I have this wrong, and I would want to know.
What to actually do about it
Watching signals is for analysts. If you run a budget, the question is operational. Almost nobody has sorted their workloads by whether they need frontier capability. It gets discussed, then it sits on the list.
A rough method: pull last month’s token spend by application. Flag anything doing classification, extraction, summarization, translation, first drafts or routine code completion. That bucket is usually 60 to 80 percent of volume and rarely needs a frontier model. Run a week of parallel traffic against a cheap open model, measure task completion rather than benchmark scores, and keep the expensive model for what actually fails.
Flag these first: Classification, Extraction, Summarization, Translation, First drafts, Routine code completion.
Usually 60 to 80 percent of volume, and rarely a frontier job.
At a 35-to-1 input price gap the arithmetic is blunt. A team spending $40,000 a month with 70 percent of volume in that bucket carries about $28,000 where the cheap option runs closer to $1,000. Allow for higher retry rates and some workloads bouncing back, and the saving is still not marginal.
Run it on your own bill, at a 35-to-1 input price gap:
| Input | Value |
|---|---|
| Monthly token spend | $40,000 |
| Share in the commodity bucket | 70% |
| Output | Value |
|---|---|
| Carried in that bucket today | $28,000 |
| The cheap option, rounded up for retries | $1,000 |
| Savings a month, before workloads bounce back | $27,000 |
Coding tool vendors, support platforms and translation shops moved first, because their margins are directly exposed to inference cost. Regulated industries have not, and probably will not.
One thing to know before you route anything: because the weights are open, Western hosts serve the same Chinese models with no data touching China, at roughly double first-party pricing and still far below frontier rates. For most companies that is the sensible version of this trade.
Western hosts serving open-weight Chinese models charge roughly double the first-party price, with no data touching China.
— OpenRouter, June 2026






