Time for Open Models to Catch Up Is Halving Each Generation

A report published by SemiAnalysis on August 21 plots the gap between open models and the closed frontier over the past three and a half years as a single curve. Its rule is blunt: every time open source closes a generation's gap, it takes half as long as the generation before.

The authors split the history into three phases. In the early scaling phase, measured against GPT-3.5 Turbo, the open side started with Llama-2-70B and didn't close the gap until Llama-3.1-405B arrived in July 2024 — a span of more than a year. The reasoning phase moved noticeably faster: the 12.1-point gap set by o1-preview was erased by R1-0528 in 8.5 months. In the agentic phase, Kimi K2.6 scored 56.3 to pass Opus 4.5 in 4.8 months, and GLM-5.2 scored 72.4 to pass GPT-5.2 in 6 months.

More than a year, then 8.5 months, then five to six months — lined up together, the three windows trace a halving slope.

Why catching up keeps accelerating

The report attributes the acceleration to several things happening at once: open weights let capability be lifted directly, distillation lets scores be transferred cheaply, and even the reinforcement-learning environments behind frontier gains are being publicly replicated. The authors put it almost like a verdict — nothing stays secret forever, especially given how much work distillation is doing.

Running the curve forward a bit (an extrapolation the report itself doesn't make): if the halving pattern holds for one more generation, the next catch-up window would shrink to two or three months — effectively same-generation, same-period releases. For closed labs, that pressure will hit pricing before it hits leaderboards. How long a model can command a premium depends on how long its lead lasts, and as that lead compresses from a year to a quarter, the window for rethinking pricing strategy compresses just as fast.

The commoditization question

The report doesn't dodge the question the market is most worried about: if open source can stay close enough to the frontier at a fraction of the cost, does the model layer get commoditized? The authors write that such an outcome would clearly be disastrous for frontier labs' margins.

They then lay out their reasoning. Open models such as GLM 5.3 and Kimi K3 are already, in the report's own phrasing, "genuinely capable" of handling much of the coding and agentic work that has pushed Anthropic's annualized revenue past $65 billion — meaning open source can genuinely do comparable work, not merely post comparable scores.

The demand side is spelled out too: the report says Fireworks alone now processes more than 40 trillion tokens a day, double the OpenAI API's volume as of late March. The business of hosting open-weight inference is growing, which is both evidence that open models' capability is real and one of the channels through which margins at the model layer get squeezed.

The authors hit the brakes themselves

Even so, the conclusion doesn't turn bearish on the frontier. The report is upfront that benchmarks are only part of the picture: Kimi K3 scores higher, yet the authors say they still reach for Fable 5 for their day-to-day work. They give two reasons — the productized experience gap that scores don't capture, and the fact that scores themselves can be inflated by narrowly targeted reinforcement-learning environments, leaving a systematic gap between leaderboard position and how a model actually feels to use. They land on a fairly measured line: things are "not as bearish for frontier models as you might initially think."

For readers in China, the report carries another layer of meaning. In the third and final stage of the curve, the models doing the catching up are almost entirely open-weight releases from Chinese teams — Kimi, GLM, and the next generation the report names. The very releases overseas analysts are citing as proof that "open source is accelerating" are, largely, this cohort's own release cadence. Seen from the other side, the same data shows that Chinese vendors have made open weights a primary battleground rather than a fallback option outside of closed models.

As for what's left for the closed-source side, the report's answer points to productization, distribution, and enterprise delivery. None of that shows up on a benchmark table, but it's what decides who ends up holding numbers like $65 billion.

Sources: SemiAnalysis, "Are Open Models Catching Up?"; editorial compilation by CocoLoop; figures cross-checked against the original report for catch-up timelines by era, the Kimi K2.6 and GLM-5.2 scoring basis, and Fireworks' daily token volume.