Kimi K3 had just put 2.8 trillion parameters and a 1 million-token context window into Moonshot AI's flagship story. Then a very different scorecard arrived.
On July 23, the UK AI Security Institute and the US Center for AI Standards and Innovation published a joint assessment of Kimi K3 on offensive cyber tasks. The result is blunt: Kimi K3 beats Zhipu's GLM-5.2, but remains far behind the leading US cyber-capable frontier models.
ExploitBench shows the gap
On ExploitBench, a 41-task public benchmark for exploit development, Kimi K3 scored 32.2%. GLM-5.2 scored 24.4%, while the leading unnamed US models averaged 76.2%. Kimi K3 achieved zero arbitrary-code-execution outcomes across all 41 tasks; the US frontier group achieved 20.
One full chain still matters
The second test, The Last Ones, is a private 32-step enterprise-network attack simulation that AISI says typically takes a human expert 20 hours. Kimi K3 reached step 17 on average, while leading US models reached 28.5. It still completed the full attack chain once in 10 attempts, which is why the result is not harmless.
The practical reading is narrower than the panic cycle. Kimi K3 does not look like a peer of the strongest US cyber models, but it can still operate inside weakly defended systems. That is the risk zone enterprises need to care about.
General scores are not cyber scores
The Decoder ties the result to the broader distillation debate. If a model learns heavily from public frontier-model interfaces, advanced cyber behavior may be missing because those public interfaces usually block high-risk exploit requests. General benchmarks, coding demos and cyber capability therefore need to be read separately.
Moonshot has not issued a public response to this assessment. The next checks are concrete: whether Kimi K3 gets a fuller cyber safety card, whether high-risk request refusal changes, and whether independent teams can reproduce these results.
Sources: UK AISI/CAISI joint assessment, US NIST notice, The Decoder, CocoLoop, South China Morning Post, Kimi K3 official technical blog; verified ExploitBench task count, 32.2%/24.4%/76.2% scores, arbitrary-code-execution counts, The Last Ones 32-step setup, one complete run in 10 attempts, and Kimi K3 parameter/context claims.