Charted: Chinese Models Now Process More AI Tokens Than American Ones (2022–June 2026)
Loading interactive charts…
The blind spot in every Western token chart
Tokens are the atomic unit of AI economics — they drive inference cost, API pricing, and data centre power draw. Almost no lab publishes a clean "tokens processed" figure, so most public estimates are built from whatever US companies happen to mention on stage.
That produces a badly distorted picture. An earlier version of this dataset tracked eight American and Canadian brands and put global volume near 3.8 quadrillion tokens a month. It was missing the single largest token processor on earth, and eight more Chinese providers behind it.
The rebuilt dashboard above tracks 836 monthly records across 19 providers — ten Chinese, seven American, plus Mistral and Cohere — from November 2022 through June 2026. Rows sitting on a published figure are flagged disclosed; the rest are interpolated between anchors.
The numbers do not reconcile, and that is the story
China's National Data Administration says the entire country processes 140 trillion tokens a day (March 2026). ByteDance says one of its model families does 180 trillion. Both cannot be true on a common basis.
At Volcano Engine's FORCE conference in June 2026, president Tan Dai disclosed that the Doubao model family processes over 180 trillion tokens per day — roughly 5.4 quadrillion per month, or 1,500x its May 2024 launch volume. ByteDance has published a figure at nearly every FORCE conference since December 2024, which makes it the fullest public series of any provider on earth:
| Date | Tokens per day | Monthly equivalent |
|---|---|---|
| May 2024 (launch) | 120B | 3.6T |
| Dec 2024 | 4T | 120T |
| Mar 2025 | 12.7T | 381T |
| Sep 2025 | 30T | 900T |
| Dec 2025 | 50T | 1,500T |
| Mar 2026 | 120T | 3,600T |
| Jun 2026 | 180T | 5,400T |
Tan Dai also confirmed the crucial caveat himself: the figure "includes both internal and external use, as well as usage within the Doubao App itself." It is a company-wide inference meter, not a sales figure.
That makes it comparable to exactly one Western number. Google disclosed 3.2 quadrillion tokens per month at I/O 2026 — also all-surfaces, spanning Search, YouTube, Workspace and Gemini — up 7x from 480 trillion a year earlier. OpenAI and Anthropic disclose nothing at all; OpenAI's only public throughput figure is 15 billion API tokens per minute as of March 2026.
Doubao's claim exceeding Google's, on a far smaller user base, should not be read as a clean result. ByteDance credits AI video generation for much of the growth, and video tokenises into enormous counts. The company has not published its counting rules.
What the market actually pays for
Set the headline claims aside and look at what customers buy. IDC measures public-cloud model services sold to external customers only, explicitly excluding first-party calls from Douyin, the Doubao app and Jimeng. On that basis China ran 1,944 trillion tokens across all of 2025 — about 5.3T/day, against headline claims 10 to 20 times larger.
The revenue attached to that volume is the most striking number in the dataset. China's entire public-cloud model market earned RMB 3.07 billion (~$430 million) in 2025. That works out to roughly 22 cents per million tokens, blended nationwide. Vercel's gateway data shows the same divergence from the demand side: DeepSeek took 17% of routed token volume while accounting for about 1% of spend.
A useful rule of thumb: divide Chinese headline figures by five to ten to approximate commercially-served tokens.
February 2026: the crossover
The cleanest evidence comes from OpenRouter, which routes third-party developer traffic across labs and publishes volume by model. What makes it credible is its audience: 47% of its developers are American and only 6% are Chinese, so Chinese models winning there reflects global demand, not domestic accounting.
In the week of 9–15 February 2026, Chinese models processed 4.12 trillion tokens against 2.94 trillion for American models — the first time Chinese models led. They have held the lead since. By mid-2026 Chinese-origin models were 46.4% of all routed volume against 35.7% for American ones. Anthropic's share halved from 29.1% to 13.3% in twelve months; Meta's Llama fell below 1%.
Beware the higher numbers in circulation. Widely-quoted figures of 61% are single-week, top-ten-only measurements. The defensible all-model figure is 46%.
Price is the mechanism — but it is now reversing
Developers switched on cost per unit of work, not benchmark scores:
| Model | Lab | Origin | $/M in | $/M out |
|---|---|---|---|---|
| GPT-5.5 | OpenAI | US | $5.00 | — |
| Qwen3.7-Max | Alibaba | China | $2.50 | $7.50 |
| Kimi K2.6 | Moonshot | China | ~$0.90 | ~$3.75 |
| Mistral Large 3 | Mistral | Europe | $0.50 | $1.50 |
| DeepSeek V4 Pro | DeepSeek | China | $0.435 | $0.87 |
| Tencent Hy3 | Tencent | China | ~$0.14 | ~$0.56 |
| DeepSeek V4 Flash | DeepSeek | China | $0.14 | $0.28 |
| MiMo-V2.5 | Xiaomi | China | $0.14 | $0.28 |
DeepSeek V4 Pro's output price sits roughly 34x below GPT-5.5. Cache-hit discounts go further still — DeepSeek charges $0.0028 per million on cached input, optimising hard for repetitive agentic prefill.
But the "China is cheap" story is inverting, and almost nobody has written it. Zhipu raised API prices 83% cumulatively and still reports customers queuing. Alibaba told analysts its "ability to supply this demand is not able to keep up" and expects to raise prices while costs fall. Kimi K3 launched in July 2026 at premium pricing rather than undercutting. Alibaba Cloud raised AI compute prices up to 34% in March 2026. The 2024–25 price war is over; the binding constraint has moved from demand to supply.
Two structural factors still amplify Chinese volume. Open weights mean DeepSeek, Qwen and GLM run on hardware their authors do not own and cannot meter — Qwen passed 1 billion cumulative Hugging Face downloads in January 2026, overtaking Llama, with 200,000+ derivative models. Every Chinese token figure is therefore a floor. And agentic workloads multiply consumption: the open-source agent framework OpenClaw, which spread across every Chinese cloud in weeks from February 2026, can burn 200,000 tokens in a single session. Coding rose from 11% of OpenRouter tokens in early 2025 to over 50% by year-end.
The official Chinese numbers
China is the only country publishing an official national series:
- Early 2024: ~100 billion tokens/day
- June 2025: over 30 trillion/day
- End 2025: 100 trillion/day
- March 2026: over 140 trillion/day
Frost & Sullivan's enterprise study — which includes private and on-premises deployment — found corporate calls rising from 10.2T/day in H1 2025 to 37.0T/day in H2 2025, with Alibaba Qwen at 32.1% share, Doubao at 21.3% and DeepSeek at 18.4%.
Note that IDC and Frost & Sullivan rank different leaders and both are correct. IDC covers external public cloud only and puts Volcano Engine first at 49.5%. Frost & Sullivan includes private deployment, where open-weight Qwen dominates. The two share tables must never be combined.
Read these numbers carefully
- Headline figures are company-wide meters, not sales. ByteDance's own executives say so. Only ByteDance and Google publish on a comparable all-surfaces basis.
- Summing providers overshoots the national total by about 64%open weights get re-served by third-party clouds and counted twice.
- Chinese tokenisation is a red herring. Measured on identical meaning, Chinese needs only 8–12% more tokens on Chinese-native tokenizers. It explains almost none of the volume gap.
- The unit itself is unstable. Reasoning models emit far more tokens per answer than 2023-era chat, and agentic requests use roughly 15x more tokens than human chat. A meaningful part of the "1,000x growth" is redefinition, not usage.
- Two crossover dates, both real. DeepSeek's R1 moment briefly put Chinese providers ahead in early 2025 before US providers regained the lead; the durable crossover in this dataset is January 2026. On OpenRouter's routed traffic it is February 2026.
What to watch
- Whether US labs start disclosingGoogle publishes throughput every quarter; OpenAI and Anthropic still do not
- Huawei's memory supplyDeepSeek has said further price cuts depend on Ascend 950 supernodes scaling. The real ceiling is CXMT's HBM stacking yield, not logic dies, which makes Chinese token prices a function of Chinese memory yield
- Whether rising prices slow adoptionthe switching argument weakens if the 30x spread compresses
- DeepSeek's shapeits curve is non-monotonic, and charts drawing it as smooth exponential growth are wrong
- A standardised metrican audited "tokens served" definition would resolve most of the ambiguity here