China's Daily AI Token Calls Surge to 500 Trillion in Three Months

China Central Television (CCTV) reported today that as of June this year, China's daily average token calls surpassed 500 trillion. The same metric was officially reported just three months earlier, when it stood at 140 trillion.

Retrace the line and it gets clearer. "Token" was formally translated into Chinese as "词元" (ciyuan) back in March, when the same briefing put daily average calls at more than 140 trillion — up more than a thousandfold from 100 billion in early 2024, and up roughly 40% from 100 trillion at the end of 2025. Extrapolating that pace, the increase over this second quarter's three months was several times larger than the previous quarter's — a rough calculation puts it near 3.6x, marking a clear inflection point in the second quarter.

Where the Curve Bends: Inference

According to CCTV Finance, the center of competition has shifted from raw "model IQ" to agent deployment and ecosystem building. That shift has a very concrete meaning for compute budgets: a chat-style query is a single question and answer, using a few hundred to a few thousand tokens before it's done. An agent taking on a task, by contrast, repeatedly retrieves information, reads context, calls tools, and runs multiple rounds of feedback — token consumption for the same task can differ by one or two orders of magnitude.

Liu Feng, general manager of Tencent's intelligent industry group, offered a number in the report that makes for a direct comparison: in its first week live, the official version of Hunyuan 3 saw token call volume increase 68-fold over the previous generation, Hunyuan 2. A single generational upgrade producing a nearly seventy-fold jump in call volume goes well beyond what "more users" can explain — it looks more like the entire consumption structure of a single task has changed.

That also explains why companies are expanding inference clusters rather than fixating only on training chips. Training is a one-time investment that gets amortized once a model finishes training; inference is an ongoing cost that grows with usage — and agents happen to be the kind of product that makes heavy usage the default.

Release Cycles Compressed to Four to Six Weeks

Feng Wen, chief architect at StepFun's (稀宇科技) open platform, pointed to a shift: in the second half of 2025, vendors were still targeting a new model release once every three months; now they need to ship a new model every four to six weeks.

That squeezes compute from both directions. On the training side, cluster scheduling is compressed — a cluster barely catches its breath before the next round begins. On the inference side, several still-active model versions need to be served at once: a new model going live doesn't mean the old version disappears immediately, and older versions often stay live on the API for months afterward, so serving capacity has to be provisioned for parallel support.

The faster pace has a knock-on effect too: the window for amortizing a single model's training cost has shortened. Domestic models have cut prices repeatedly this year, and the price war is closely tied to that narrowing window — it makes more sense to maximize usage while a model is still in active service than to sell it slowly at a high price.

AI Data Centers Are Chasing Long-Term Contracts

Zheng Zihao, general manager of the AI computing center at Envision Group (远景科技集团), described the buildout of AI data centers as booming, citing long-term, stable order commitments from customers as the reason.

Those five words — "long-term stable orders" — are the key. Data centers are capital-heavy assets; it takes one to two years to go from securing land to switching a facility on, and investors' biggest fear is building it out only to have it sit idle. With long-term contracts as a backstop, builders are willing to design for full capacity from the start — power supply, liquid cooling, and rack density can all be built out properly the first time. The domestic "token factories" and integrated AI computing hubs now underway in multiple regions follow the same logic: lock in demand first, then build capacity, pushing the domestic compute supply chain from product validation to volume delivery.

What's Still Missing Beyond the Number

The 500 trillion figure still lacks a horizontal benchmark. There's no publicly available global figure using the same methodology, so "first tier" remains, for now, a qualitative claim rather than something backed by a ratio.

Spread across the population, 500 trillion tokens works out to roughly 350,000 tokens per person per day across China's 1.4 billion people, by rough calculation. That number isn't really meaningful on its own — the vast majority of calls come from enterprise systems and backend agents, unrelated to how much any individual is typing. Its real use is as a reminder: tokens have shifted from being a technical metric to an industrial one. What gets compared going forward will be the electricity bill, the GPU-hours, and the revenue that each trillion tokens brings back.

Sources: CCTV Finance, IT Home, CocoLoop, National Data Administration public briefings; call-volume figures were checked line by line against the previous official briefing, and the per-capita figure is the editor's rough estimate.