Skip to content

Intelligence Layer

One account, one ledger, one agent layer. The modules are what you switch on.

Platform overview
The AI Intelligence Layer

One intelligence layer under every module: the same data, the same agents, the same ledger, whichever function you switch on.

  • One layer, every module
  • Shared data and agents
  • Every run on one ledger
Go to The AI Intelligence Layer
Isometric stack of pale grey slabs wired together on white, with one layer filled deep cyanStart hereThe architecture is the argumentWhy adding the sixth module costs a fraction of buying the first.Read the architecture
The Architecture

One account carries every brand you run. Modules are switched on per brand, and every run lands in the same ledger.

  • One account, many brands
  • Modules toggle per brand
  • A single billing ledger
Go to The Architecture
Isometric stack of pale grey slabs wired together on white, with one layer filled deep cyanStart hereThe architecture is the argumentWhy adding the sixth module costs a fraction of buying the first.Read the architecture
The agent layer

Three kinds of agent. Watchers notice, drafters produce, and Ask Intelligence answers from your own numbers.

  • Watchers run on a schedule
  • Drafters wait for sign-off
  • Answers carry their source
Go to The agent layer
Isometric stack of pale grey slabs wired together on white, with one layer filled deep cyanStart hereThe architecture is the argumentWhy adding the sixth module costs a fraction of buying the first.Read the architecture
Security

Isolated tenancy, roles that mean something, and a plain answer about what leaves your account.

  • Per-tenant isolation
  • Roles down to the module
  • Stated data boundaries
Go to Security
Isometric stack of pale grey slabs wired together on white, with one layer filled deep cyanStart hereThe architecture is the argumentWhy adding the sixth module costs a fraction of buying the first.Read the architecture

Seats are free — you pay for what the AI runs

Book a demoOpen your account

Capabilities

The AI you work with directly, and the modules it already runs. Take one, or take them all.

See all modules

AI Assistance

Work with the AI directly

Ask Intelligence

Ask a business question in plain words. It reads your own data through read-only tools scoped to the brand, and answers with the real figure or says the data is not there.

  • Fifteen read-only tools
  • Scope enforced in code
  • Never invents a number
Go to Ask Intelligence
Playground

A direct chat with the models your administrator allows. No brand data: the model sees the conversation, your files and your instructions.

  • Compare three models side by side
  • Files in, documents out
  • The price under every answer
Go to Playground
Agents

Three kinds of agent do the standing work: watchers notice, drafters produce, and nothing ships until someone with the role signs.

  • Watchers run on a schedule
  • Drafters wait for sign-off
  • Every run lands on the ledger
Go to Agents
Skills

How a brief is built, how a campaign is structured, what a decent article looks like — the craft of the people who built this, shipped as the default setting.

  • Senior practice built in
  • Defaults you adjust
  • Nothing to write from scratch
Go to Skills

Modules & Features

The functions it runs

Business Intelligence

Ask your business a question in plain words and get the real number back, with the source that produced it.

  • Plain-language questions
  • Every figure sourced
  • Board-ready reporting
Go to Business Intelligence
Project Management

The board where an insight becomes a card with an owner, instead of a dead report nobody actions.

  • Kanban with a timeline
  • One click, insight to task
  • Owner and date attached
Go to Project Management
Sales

Pipeline that drafts its own paperwork, and flags the deals that went quiet before you notice.

  • Quotes off the deal record
  • Quiet deals surfaced
  • Revenue projected forward
Go to Sales
Brand Marketing

The content engine and the Creative Studio behind the brand: articles buyers actually search for, and the images and video that carry them.

  • Eight-stage content engine
  • Creative Studio built in
  • Drafts wait for sign-off
Go to Brand Marketing
Social

Posts drafted from work you have already approved — an article becomes the thread, a launch becomes the post — on a calendar you sign.

  • Drafts from approved work
  • A calendar you sign
  • One voice on every network
Go to Social
Search

Watches where you rank, ties every position to the page that earned it, and says what to write next — with GEO watching the AI assistants.

  • Rank tracked continuously
  • Positions tied to pages
  • GEO polls the assistants
Go to Search
Ads

Search campaigns arrive built — ad groups, keywords, copy — priced before they run, and paused the moment they stop earning.

  • Campaigns priced first
  • Ad groups arrive built
  • Paused when they slip
Go to Ads
Customer Service

Answers from what you gave it, on your own page, and it says so plainly when it does not know.

  • Grounded in your content
  • Refuses to invent
  • Escalates to a human
Go to Customer Service
IT Support

The support desk turned inward: employees ask, it answers from your own systems and documentation, and escalates what it cannot resolve.

  • Answers from your docs
  • Tickets triaged first
  • Escalates to a human
Go to IT Support

Seats are free — you pay for what the AI runs

Book a demoOpen your account

Services

The people around the product, from the first connection to a named operator on your account.

All services
Onboarding

A guided session to connect your data and switch on the first function, then a first week planned day by day.

  • One guided session
  • First function live in a week
  • A person on the other end
Go to Onboarding
A fanned stack of white ruled sheets with the top one edged in deep cyanLarger organizationsRoll out by brand, by team, by functionOne account, one ledger, roles down to the function. Staged so procurement never has to open a file.See the group rollout
Data integration

We connect the systems you already run and prepare the data behind them, so every answer has a source.

  • Connectors built and tested
  • Metrics defined once
  • Sources reconciled
Go to Data integration
A fanned stack of white ruled sheets with the top one edged in deep cyanLarger organizationsRoll out by brand, by team, by functionOne account, one ledger, roles down to the function. Staged so procurement never has to open a file.See the group rollout
Intelligence audit

A fixed-price review of where you rank, what the AI assistants answer, and which function pays back first.

  • Four assistants polled
  • Search position by market
  • A ranked starting point
Go to Intelligence audit
A fanned stack of white ruled sheets with the top one edged in deep cyanLarger organizationsRoll out by brand, by team, by functionOne account, one ledger, roles down to the function. Staged so procurement never has to open a file.See the group rollout
Training

Sessions for the people who will approve, review and ask, so the workforce is used rather than watched.

  • By role, not by feature
  • Live on your account
  • Recorded for the next hire
Go to Training
A fanned stack of white ruled sheets with the top one edged in deep cyanLarger organizationsRoll out by brand, by team, by functionOne account, one ledger, roles down to the function. Staged so procurement never has to open a file.See the group rollout
Managed operations

Someone reviews what the watchers found, approves the drafts within your caps, and runs the weekly review.

  • Drafts approved in your name
  • Caps respected
  • Weekly review delivered
Go to Managed operations
A fanned stack of white ruled sheets with the top one edged in deep cyanLarger organizationsRoll out by brand, by team, by functionOne account, one ledger, roles down to the function. Staged so procurement never has to open a file.See the group rollout
Custom features

An agent or a function built for your process, on the same ledger and under the same approval gates.

  • Scoped before it is priced
  • Same gates, same ledger
  • Yours to keep
Go to Custom features
A fanned stack of white ruled sheets with the top one edged in deep cyanLarger organizationsRoll out by brand, by team, by functionOne account, one ledger, roles down to the function. Staged so procurement never has to open a file.See the group rollout
Dedicated hosting

A single-tenant deployment in the region you choose, with the data boundary written down and an uptime commitment.

  • Single tenant
  • Region of your choice
  • Uptime in writing
Go to Dedicated hosting
A fanned stack of white ruled sheets with the top one edged in deep cyanLarger organizationsRoll out by brand, by team, by functionOne account, one ledger, roles down to the function. Staged so procurement never has to open a file.See the group rollout
Premium Support

A named contact, committed response times, and a quarterly review of what ran, what it cost and what to change.

  • Named contact
  • Committed response times
  • Quarterly review
Go to Premium Support
A fanned stack of white ruled sheets with the top one edged in deep cyanLarger organizationsRoll out by brand, by team, by functionOne account, one ledger, roles down to the function. Staged so procurement never has to open a file.See the group rollout

Seats are free — you pay for what the AI runs

Book a demoOpen your account

About

Why this was built, who is behind it, what it has done for the companies running it, and how to run it well.

About BearingBridge
Why we built it?

What a company can do has been capped by who it could afford to hire. That cap is the thing that moved.

  • Capability, not headcount
  • Written, not benchmarked
  • No lock-in claim
Go to Why we built it?
A row of pale grey blocks climbing steeply to a single tall deep cyan block, a hairline curve tracing the riseThe short versionPunch above your headcountThe AI workforce for companies that need more capability than they can staff.Read the argument
Who is behind BearingBridge?

Who built this, what it is independent of, and why that independence is worth stating out loud.

  • Model-independent
  • No data resale
  • Named people behind it
Go to Who is behind BearingBridge?
A row of pale grey blocks climbing steeply to a single tall deep cyan block, a hairline curve tracing the riseThe short versionPunch above your headcountThe AI workforce for companies that need more capability than they can staff.Read the argument
How we run the work

The six-phase method every account is built on: a written bearing, a dated baseline, a pilot on real data, and a fix that reconciles the money.

  • A kill criterion, in writing
  • Costs watched as they run
  • An ending without us
Go to How we run the work
A row of pale grey blocks climbing steeply to a single tall deep cyan block, a hairline curve tracing the riseThe short versionPunch above your headcountThe AI workforce for companies that need more capability than they can staff.Read the argument
Case studies

Ten builds told the way they ran: the situation, what was built, what changed months later, and what got switched off.

  • Sector and function stated
  • Client-verified figures
  • The stopped work left in
Go to Case studies
A row of pale grey blocks climbing steeply to a single tall deep cyan block, a hairline curve tracing the riseThe short versionPunch above your headcountThe AI workforce for companies that need more capability than they can staff.Read the argument
FAQ

The questions that come up before every demo, answered here so the demo can be about your business.

  • Product and billing
  • Security and data
  • Answered in plain words
Go to FAQ
A row of pale grey blocks climbing steeply to a single tall deep cyan block, a hairline curve tracing the riseThe short versionPunch above your headcountThe AI workforce for companies that need more capability than they can staff.Read the argument
Testimonials

Customers on the module they actually run, quoted directly, with the module named.

  • Module named each time
  • Quoted, not paraphrased
  • Role and size given
Go to Testimonials
A row of pale grey blocks climbing steeply to a single tall deep cyan block, a hairline curve tracing the riseThe short versionPunch above your headcountThe AI workforce for companies that need more capability than they can staff.Read the argument
Changelog

What shipped, dated, newest first. The place to check whether the thing you were promised exists yet.

  • Dated entries
  • Shipped only
  • Linked to the module
Go to Changelog
A row of pale grey blocks climbing steeply to a single tall deep cyan block, a hairline curve tracing the riseThe short versionPunch above your headcountThe AI workforce for companies that need more capability than they can staff.Read the argument
Partner program

Run the platform for the companies you advise. Every client is a brand on your account, billed on its own ledger.

  • Clients as brands
  • Separate ledgers
  • Margin on every run
Go to Partner program
A row of pale grey blocks climbing steeply to a single tall deep cyan block, a hairline curve tracing the riseThe short versionPunch above your headcountThe AI workforce for companies that need more capability than they can staff.Read the argument

Seats are free — you pay for what the AI runs

Book a demoOpen your account

Pricing

Seats are free. You pay for what the AI actually runs, and you see the price first.

Full pricing
How it works

Three sentences, and that is the whole model. No tiers to decode, no per-seat arithmetic to do.

  • No seat licence
  • No annual lock-in
  • Three sentences long
Go to How it works
Ten pale meter tracks filled to different levels in deep cyan, cut by one horizontal cyan limit lineNo surprisesYou see the price before you run itThe price, the caps and the gates, all visible before anything spends.Open pricing
Seats

Give an account to everyone who needs one. The number on the invoice does not move when you do.

  • Unlimited accounts
  • Zero per-seat cost
  • Roles still enforced
Go to Seats
Ten pale meter tracks filled to different levels in deep cyan, cut by one horizontal cyan limit lineNo surprisesYou see the price before you run itThe price, the caps and the gates, all visible before anything spends.Open pricing
What things cost

Every run is priced from your wallet before or as it runs, and the machine’s own mistakes are not billed to you.

  • Quoted before it runs
  • Retries are on us
  • Itemised in the ledger
Go to What things cost
Ten pale meter tracks filled to different levels in deep cyan, cut by one horizontal cyan limit lineNo surprisesYou see the price before you run itThe price, the caps and the gates, all visible before anything spends.Open pricing
Caps and gates

A hard cap per module, and an approval gate in front of anything that publishes or spends.

  • Hard cap per module
  • Approval before publish
  • A zero balance stops the AI
Go to Caps and gates
Ten pale meter tracks filled to different levels in deep cyan, cut by one horizontal cyan limit lineNo surprisesYou see the price before you run itThe price, the caps and the gates, all visible before anything spends.Open pricing
Questions

The questions people actually ask about the bill, answered on the page rather than in a call.

  • Overage answered
  • Cancellation answered
  • Migration answered
Go to Questions
Ten pale meter tracks filled to different levels in deep cyan, cut by one horizontal cyan limit lineNo surprisesYou see the price before you run itThe price, the caps and the gates, all visible before anything spends.Open pricing

Seats are free — you pay for what the AI runs

Book a demoOpen your account

Insights

Three collections, one standard: long-form, sourced, dated, and never behind a form.

All insights
AI Trends

What is moving under the industry: model economics, the Chinese price tier, and how buyers now ask assistants instead of searching.

  • The model market, read closely
  • AI search and citations
  • Every claim sourced and dated
Go to AI Trends
A fanned stack of white ruled sheets with the top one edged in deep cyanThe standardLong-form, sourced, never gatedEach piece answers the question in its title completely, carries its date, and sits behind no form.Open the library
Best Practices — AI Guide

The working guide: when to buy and when to build, when an agent is the wrong tool, and what survives contact with production.

  • Build-or-buy, decided
  • Architectures that ship
  • Prompts that do real work
Go to Best Practices — AI Guide
A fanned stack of white ruled sheets with the top one edged in deep cyanThe standardLong-form, sourced, never gatedEach piece answers the question in its title completely, carries its date, and sits behind no form.Open the library
CEO's Opinion

Signed columns from the person running the company: what building agents across six functions actually shows, ahead of the industry line.

  • Signed, never ghostwritten
  • From live builds, not decks
  • Positions, not press releases
Go to CEO's Opinion
A fanned stack of white ruled sheets with the top one edged in deep cyanThe standardLong-form, sourced, never gatedEach piece answers the question in its title completely, carries its date, and sits behind no form.Open the library

Seats are free — you pay for what the AI runs

Book a demoOpen your account

Insights/AI Trends/East-West model notes

China wants to industrialize the token.

Five assumptions hold the export story up. The price gap survives. The chips, the idle racks and the revenue do not.

140T

daily tokens, state-reported

36.8%

effective compute usage

46%

of revenue, 12% of volume

BearingBridgeJuly 20269 min read

A technician walks a service aisle in a computing hall, half the racks unlit, network cabling running loose along a scuffed concrete floor under one warm service lamp
On this sheet

At WAIC in Shanghai this month, across 1,100 exhibitors and more than 140 forums, the phrase that kept coming up was not AGI. It was Token工厂, token factory.

There is nothing subtle about the strategy behind it. Take a unit of AI output and standardize it. Industrialize the production. Drive the cost down until nobody else wants to compete, then export the surplus. It is the solar and battery playbook, applied to inference.

I run a cross-border business out of Shanghai, and before that spent seven years running a commerce agency here on a P&L that lived or died on Chinese platform economics. That makes me skeptical in two directions: of Western coverage that waves off Chinese infrastructure, and of Chinese coverage that treats a pilot as a finished export industry.

To be clear up front, none of this argues Chinese models are weak. Stanford’s AI Index put the top US model 2.7 percent ahead of the top Chinese model in March. The capability question is close to settled. What follows is about infrastructure and economics.

Five assumptions hold the strategy up. One stands. Four are shaky. If you only want the practical part, skip to the last section.

The assumption ledger — open a row to see what the available data does to it (1/5 stands)

Assumption Meter Verdict Detail
The export loop works 22% One pilot Shantou closed the loop in April: data in, inference domestic, answer out by API. No volume and no revenue have been published, and the city was picked for a cable landing station most of the country does not have.
The price gap is real 92% Stands Open-weight Chinese models run 60 to 90 percent below leading Western pricing, and developers have already routed accordingly. This is the strongest leg the strategy has.
Domestic silicon can carry it 30% Not yet Huawei’s 910C delivers roughly 60 percent of H100 inference performance, and most of DeepSeek’s tokens are still inferenced on Western hardware.
The factories are running 28% Mostly idle Effective usage of 36.8 percent in 2025, rack utilization of 20 to 30 percent at many centers, against leading platforms running at 90 to 95 percent. A coordination problem, not a supply glut.
Volume converts to revenue 18% Inverted Anthropic holds around 12 percent of token volume on OpenRouter and captures close to 46 percent of the revenue. China owns the commodity lane; the profit sits in the other one.

Legend: Survives the data · Running ahead of the evidence.

36Kr’s WAIC recap called it the first year the token economy became a core conference topic.

— 36氪, July 2026

Make inference a metered commodity

In March, China’s National Data Administration fixed the official Chinese term for token as 词元, cí yuán, and called it a settlement unit linking technical supply to commercial demand. Beijing now tracks daily token volume and reports it at policy events the way it reports steel output.

The headline number is a thousandfold rise in two years. About 100 billion daily calls in early 2024, more than 140 trillion by March 2026.

  • 100B daily calls, early 2024
  • 140T daily, March 2026 — a 1,000× rise
  • State-reported. Unaudited. Counts consumption, not value.

Treat that with some care. It is state-reported, nobody audits it, and it counts consumption instead of value. That is not a China problem specifically. No government or vendor anywhere publishes an audited token number. An agent stuck in a bad loop overnight burns an absurd quantity and produces nothing.

By end-2025 China had built more than 100,000 high-quality datasets totaling over 890 petabytes, the supply side of the same policy push.

— 新华社, March 24, 2026

The mechanism is real, but it is one pilot

In April, Shantou in Guangdong closed a loop under a policy called 来数加工, inbound data processing. Overseas data enters a digital bonded zone. Domestic compute runs the inference. The answer goes back out by API. Power, compute and revenue all stay in China.

  1. Overseas data in — enters a digital bonded zone
  2. Domestic compute — runs the inference
  3. Answer out by API — only the output crosses back

Power stays. Compute stays. Revenue stays.

Chinese state media summed it up with a line that traveled widely: the salt can travel the world, but the salt fields must stay home. Good line. I have caught myself repeating it. It is also promotional framing from the same system that publishes the consumption figures.

What has not been published is volume. One city, one pilot, no disclosed revenue. Shantou was picked for reasons that predate AI entirely. It holds one of three mainland submarine cable landing stations, and latency to Singapore runs about 32.7 milliseconds. Most of the country has neither. That part tends to get left out.

On the record:

  • 1 of 3 mainland submarine cable landing stations
  • 32.7 ms latency to Singapore

Not published:

  • n/a export volume
  • n/a disclosed revenue

Chinese coverage frames token exports as an impossible triangle of latency, data compliance and very cheap green power.

— 澎湃新闻 and 同花顺财经, 2026

It is real, and developers have already voted

Nobody argues with this one. It is the strongest leg the strategy has.

DeepSeek V4 Flash costs $0.14 per million input tokens. GPT-5.5 costs $5.00. Open-weight Chinese models run 60 to 90 percent below leading Western pricing. On OpenRouter, which routes calls to whichever provider wins on price, Chinese models went from under 10 percent of usage in early 2025 to a record 58 percent of tokens processed by US firms in July.

Model Input, per 1M tokens
DeepSeek V4 Flash $0.14
GPT-5.5 $5.00

Open weights run 60 to 90 percent below leading Western pricing.

Chinese model share on OpenRouter (of tokens processed by US firms): under 10% in early 2025, rising to 58% in July 2026, a record.

In agent workloads the gap compounds, because one task can fan out across dozens of parallel sub-agent calls.

OpenRouter weekly volume grew from roughly 5 trillion tokens in April 2025 to over 20 trillion a year later.

— CNBC, July 2026

Where it breaks first: the compute underneath

This is the gap most coverage understates, and it is the one that bothers me. You cannot industrialize output on equipment you cannot get enough of.

Huawei’s 910C delivers roughly 60 percent of H100 inference performance, by DeepSeek’s own published assessment. The Council on Foreign Relations puts the best US chips at about five times more powerful than Huawei’s best today, widening toward 2027. Then there is the SemiAnalysis detail, which matters more than any chip spec: most of DeepSeek’s tokens are still inferenced on Western hardware. The flagship of Chinese open-weight AI is not running the Chinese stack at scale.

Comparison Value Note
Huawei 910C vs H100 60% of H100 inference performance
Best US chip vs best Huawei more powerful today, widening toward 2027
CloudMatrix 384 vs NVL72 4.1× the power, for about 1.7× the compute

The flagship of Chinese open-weight AI is not running the Chinese stack at scale.

The energy argument, which usually gets made next, has the same problem. China generates more than twice the electricity of the United States and added 543 gigawatts in 2024 alone. But a four-fold power penalty on domestic accelerators eats much of a two-fold price advantage. Cheap power ends up subsidizing inefficient silicon rather than producing cheaper tokens.

The advantage: 2× the electricity of the United States, plus 543 gigawatts added in 2024 alone — eaten by the penalty: 4× the power draw on domestic accelerators, against a two-fold price advantage.

Chinese firms are not sitting still. Zhipu trains and serves on domestic accelerators. Alibaba is building data centers across South Korea, Malaysia, Thailand and Mexico, renting foreign capacity to sidestep the silicon problem. Both are rational moves, and neither closes the gap this year.

Huawei’s CloudMatrix 384 draws roughly 4.1 times the power of Nvidia’s NVL72 for about 1.7 times the compute.

— SemiconductorX, 2026

A lot of the capacity is sitting idle

The factory framing implies the plants are running. Chinese domestic reporting suggests many are not.

Tencent Cloud puts average GPU utilization at China’s intelligent computing centers below 30 percent. CAICT data cited by Huxiu shows an effective usage rate of 36.8 percent in 2025, meaning roughly 70 yuan of every 100 invested is idling. Securities Times reported rack utilization of 20 to 30 percent at many centers.

  • Below 30% — average GPU utilization at intelligent computing centers (Tencent Cloud)
  • 36.8% — effective usage rate in 2025 (CAICT, via Huxiu)
  • 20 to 30% — rack utilization at many centers (Securities Times)

Calling this oversupply misses it. The market has split, and I have watched the same pattern in Chinese logistics: healthy national totals sitting on stranded regional capacity. Leading platforms run at 90 to 95 percent with order books into 2028. The high-quality compute is genuinely short, while the low-quality stuff cannot find a customer at any price. That is a coordination problem, and those take years to clear.

  • High-quality compute — leading platforms run at 90 to 95 percent with order books into 2028. Genuinely short. (order books into 2028)
  • Stranded capacity — the low-quality stuff cannot find a customer at any price. A coordination problem, and those take years to clear. (no customer at any price)

A thousand-card computing center costs over 30 million yuan a year to operate, which is why low utilization turns into losses fast.

— 虎嗅 and 观察者网, July 2026

Volume is not revenue, and that is the gap that matters

If you read one number here, make it this one. Anthropic holds around 12 percent of token volume on OpenRouter and captures close to 46 percent of the revenue.

Metric Anthropic share, on OpenRouter
Share of token volume 12%
Share of revenue 46%

The lanes are priced differently.

The market is splitting into a commodity lane and a premium lane. China owns the commodity lane. The profit sits in the other one.

The economics are moving the wrong way for everyone. Huxiu documents Alibaba Cloud and Baidu AI Cloud raising prices by up to 34 percent this year. Agents consume 100 to 1,000 times what a chatbot does. Unit prices keep falling while invoices keep growing, and nobody has an accepted way to link tokens consumed to work completed.

  • Up to 34% — price rises at Alibaba Cloud and Baidu AI Cloud this year
  • 100 to 1,000× — what an agent consumes against a chatbot
  • 85% — of enterprise AI budgets now going to inference

Inference now accounts for roughly 85 percent of enterprise AI budgets, against near-zero marginal cost in traditional software.

— 腾讯新闻, June 2026

The one no price cut touches

Under China’s 2017 National Intelligence Law, companies must cooperate with state intelligence work. That is a jurisdictional fact. It says nothing about how any particular company behaves, and it is still enough to keep regulated finance, healthcare and government workloads with Western providers.

Washington is escalating separately. The Treasury Secretary raised the prospect of sanctions in July over alleged distillation of US models, days ahead of a September bilateral AI dialogue. If part of the Chinese cost advantage comes from not paying frontier training costs, it is less durable than it looks. Nobody has proven that case. It is not trivial either.

The practical response The practical response is unglamorous. If you route production traffic through Chinese-origin endpoints, keep a tested fallback configured. Model availability became a policy variable this year, and it cuts both ways.

Large enterprises continue buying mostly from Anthropic, OpenAI, Azure and Google Cloud.

— Baiguan, July 2026

What would change my mind

I am not predicting failure. The export story is running ahead of the evidence, which is different. Three signals would move me: Shantou publishing real export volume and revenue, Chinese computing center utilization crossing 50 percent, and a regulated Western enterprise putting a production workload on a China-hosted endpoint.

  • Shantou publishes real export volume and revenue.
  • Chinese computing center utilization crosses 50 percent.
  • A regulated Western enterprise puts a production workload on a China-hosted endpoint.

Any of those would tell you more than another month of routing charts. If all three move, I have this wrong, and I would want to know.

What to actually do about it

Watching signals is for analysts. If you run a budget, the question is operational. Almost nobody has sorted their workloads by whether they need frontier capability. It gets discussed, then it sits on the list.

A rough method: pull last month’s token spend by application. Flag anything doing classification, extraction, summarization, translation, first drafts or routine code completion. That bucket is usually 60 to 80 percent of volume and rarely needs a frontier model. Run a week of parallel traffic against a cheap open model, measure task completion rather than benchmark scores, and keep the expensive model for what actually fails.

Flag these first: Classification, Extraction, Summarization, Translation, First drafts, Routine code completion.

Usually 60 to 80 percent of volume, and rarely a frontier job.

At a 35-to-1 input price gap the arithmetic is blunt. A team spending $40,000 a month with 70 percent of volume in that bucket carries about $28,000 where the cheap option runs closer to $1,000. Allow for higher retry rates and some workloads bouncing back, and the saving is still not marginal.

Run it on your own bill, at a 35-to-1 input price gap:

Input Value
Monthly token spend $40,000
Share in the commodity bucket 70%
Output Value
Carried in that bucket today $28,000
The cheap option, rounded up for retries $1,000
Savings a month, before workloads bounce back $27,000

Coding tool vendors, support platforms and translation shops moved first, because their margins are directly exposed to inference cost. Regulated industries have not, and probably will not.

One thing to know before you route anything: because the weights are open, Western hosts serve the same Chinese models with no data touching China, at roughly double first-party pricing and still far below frontier rates. For most companies that is the sensible version of this trade.

Western hosts serving open-weight Chinese models charge roughly double the first-party price, with no data touching China.

— OpenRouter, June 2026

Start

Reading about it is the slow way.

Open an account and run one module on one brand, or book a demo and see the platform on your own data.

No subscription. No seat fee. No card required to request an account.