On the first day GLM-5.3-Flash ended its free run and switched to paid pricing, DeepSeek reclaimed the top spot in OpenCode's daily usage rankings. The claim comes from OpenCode's own account, which named one more model in the same post: Muse Spark, sitting further down the table, was called a quiet dark horse.
Making sense of the shift means first looking at where GLM stood before.
What the free period did to the leaderboard
On August 26, Zhipu AI officially launched GLM-5.3-Flash and confirmed it as the model behind ox-alpha, which had been running anonymously on OpenCode and OpenRouter. During that anonymous testing period it was free for developers, and it quickly became the most-used model of the week.
OpenCode's public cumulative leaderboard, covering the window from July 4 to August 28, shows the size of the shift the free offer caused:
| Model | Vendor | Cumulative tokens | Change |
|---|---|---|---|
| glm-5.3-flash | Zhipu AI | 41T | +881% |
| deepseek-v4-flash | DeepSeek | 17T | -75% |
| mimo-v2.5 | Xiaomi | 11T | -2% |
| muse-spark-1.2-contributor | Meta | 9.5T | +180% |
| hy3 | Tencent | 3.2T | +111% |
One line up 881%, another down 75% — the two curves mirror each other, suggesting the same budget-conscious developers moved together during the weeks the free quota was available.
Working backward from the promotional price: at $0.075 per million input tokens, 41 trillion tokens would come to roughly $3.07 million. That's the revenue Zhipu AI didn't collect during this window — and effectively the price tag on the top-spot exposure it bought. Measured against what leading Chinese model makers typically spend on marketing, $3 million for a month of developer mindshare isn't expensive.
The pricing game isn't over yet
GLM-5.3-Flash's standard pricing is $0.15 per million input tokens and $0.50 per million output tokens; the promotional period cuts that in half, to $0.075 input and $0.25 output, running through 24:00 (UTC+8) on September 9. For comparison, GLM-5.3 itself is priced at $1.4 input and $4.4 output, making Flash roughly a tenth of the cost.
On specs, GLM-5.3-Flash is a natively multimodal model with 320B total parameters and 18B active parameters, a 1-million-token context window, and, according to public information, training and inference running on domestic Chinese chips. On benchmarks, Zhipu AI reports a score of 84.3 on Terminal Bench 2.1 and 63.4 on DeepSWE v1.1, with all six coding and agent evaluations surpassing the previous-generation GLM-5.2.
So the pullback doesn't amount to a verdict on the product. Losing the top daily spot on day one of paid pricing is a test of price elasticity — when the capability gap between two models falls within what developers can tolerate, the only decision variable left is unit cost and quota. DeepSeek-V4-Flash had already pushed its price down very low, so once the free quota disappeared, its original users had no reason to stay elsewhere.
How to read leaderboards like this
OpenCode ranks models by token consumption, which has the advantage of tracking real workloads — but the trade-off is that any free promotion directly inflates the denominator. The same issue shows up on OpenRouter: anonymous models, trial quotas, time-limited discounts — each one leaves a spike on the chart that has nothing to do with the model's actual capability.
For developers in China, the more useful reading is the line after paid pricing kicks in. Rankings during the free period show who was willing to spend on promotion; rankings after the first paid day start to show who actually stuck around. September 9, when the promotion ends and Flash returns to standard pricing, is the next window to watch.
Sources: OpenCode's public data page and official account, Zhipu AI's GLM-5.3-Flash release notes, CocoLoop, and public benchmark roundups; cumulative tokens and change figures are from the window shown on OpenCode's data page, and cost is a rough estimate based on the promotional input price.