Google's TPU v7 Beats GB200 by 57% Running Kimi K3

Inferact, a startup founded by core members of the vLLM team, has published a set of inference benchmarks: running Moonshot AI's Kimi K3 on 16 Google TPU v7 Ironwood chips hit a single-stream decoding speed of 709 tokens per second, compared with 452 tokens per second for a control group of 16 Nvidia GB200 chips — about 57% faster.

The benchmark report and code were released together on September 23, with the repository, inferact/tpu-megakernels, open-sourced under the Apache 2.0 license. Inferact, founded by the original vLLM team, previously raised a $150 million seed round at an $800 million valuation, led by a16z and Lightspeed.

On paper, GB200 has the edge

Start with the raw specs on both sides. Per the report's figures, a single TPU v7 chip delivers 2.31 PFLOPS of BF16 peak compute and 4.61 PFLOPS of FP8, with 206GB of HBM at 7380 GB/s of bandwidth. A single GB200 chip, by comparison, delivers 2.5 PFLOPS and 5 PFLOPS respectively, with 186GB of HBM at 8000 GB/s of bandwidth. GB200 leads on both compute and bandwidth; TPU's only edge is 20GB more memory capacity.

Kimi K3 is a 92-layer MoE model whose attention mechanism mixes Kimi Delta Attention with multi-head latent attention (MLA). The test split the model across the 16 chips using 4-way tensor parallelism and 8-way expert parallelism.

Here's the raw comparison without speculative decoding enabled:

Batch sizeTPU v7 (megakernel)GB200 (vLLM)
1249127
2392227
4515373
8865636

Units are tokens per second. At batch size 1, TPU is nearly twice as fast; by batch size 8, the lead narrows to just over 30%. With DeepSeek's DSpark speculative decoding turned on, single-stream speed becomes 709 versus 452, with per-step decoding taking about 8.5 milliseconds.

The trick: fusing 92 layers into one program

Inferact calls its approach a megakernel. In conventional inference frameworks, each layer's computation is split into several independent kernels launched one after another, with gaps between kernels — and weights can only start moving from memory to the chip once the corresponding kernel begins running.

Inferact used Pallas to write the entire 92-layer decoding step as a single program. According to the report, a megakernel erases kernel boundaries, so weight loading is no longer tied to the kernel that uses it — prefetching can start as far ahead as the chip's on-chip memory allows. During single-stream decoding, the chip spends most of its time waiting for data, and this prefetching fills exactly that gap; as batch size rises and compute takes up a larger share of the time, the payoff shrinks — which matches the trend in the table above.

A side benefit is compile time: XLA compilation used to take more than 30 minutes, while the megakernel compiles in under 90 seconds. To show the speedup didn't come at the cost of accuracy, the report includes Kimi K3's scores on TPU: 94.4% on GPQA-Diamond and 97.2% on GSM8K.

A rough read on compute efficiency

Using the single-stream speculative-decoding results for a rough calculation: TPU v7's BF16 peak is about 92% of GB200's, yet its token throughput is 1.57 times higher — working out to roughly 1.7 times the output per unit of peak compute. Normalized by bandwidth, the ratio comes out to about the same, roughly 1.7 times.

That number deserves a caveat. The GB200 side used vLLM's publicly released Kimi K3 recipe, a conventional multi-kernel approach that hasn't received the same megakernel-level optimization. The report doesn't include a GB200 result with equivalent optimization, so how much of the gap comes from the hardware versus the software investment can't be separated out for now. Nvidia hasn't responded to the numbers.

For Moonshot AI, the benchmark offers a viable deployment path for Kimi K3 on non-Nvidia hardware. Kimi K3 is a 2.8-trillion-parameter open-weight model, a scale that makes self-hosting a high bar to clear; with an open-source TPU kernel now available, renting TPUs in the cloud to run K3 becomes another option.

Sources: Inferact technical blog, CocoLoop, the inferact/tpu-megakernels code repository, QbitAI; throughput figures follow Inferact's 16-vs-16-chip test methodology, and funding details are based on public reporting.