Anthropic published a research report on September 29 evaluating the cyberattack capability of Zhipu AI's GLM-5.3. Its conclusion: this open-weight model can already build end-to-end exploits autonomously, at a level close to Anthropic's own Claude Mythos Preview, and its safety guardrails can be bypassed with fairly simple methods.
The report states directly:
"GLM-5.3 will likely give malicious actors access to capabilities that will allow them to find and exploit cyber vulnerabilities without meaningful restrictions."
What the capability tests covered
Anthropic ran three sets of tests. The first was ExploitBench, targeting vulnerabilities in the Chrome V8 engine: GLM-5.3 succeeded 50 times out of 410 attempts, a success rate of about 12%, versus 14% for Mythos Preview. The second was an internal binary-exploitation test built on the OSS-Fuzz project, where GLM-5.3 scored 4% and Mythos Preview 6%; the comparison models — Claude Opus 4.6, Kimi K3, DeepSeek V4.1-Flash, and the previous-generation GLM-5.2 — all scored 0%. The third was a set of open-ended attack tasks involving human experts.
The clearest illustration of how low the bar has dropped is a real-CVE reproduction case: with about 20 minutes of human input, GLM-5.3-Flash ran on its own for 8 hours and, at Zhipu's API pricing, produced a working exploit for roughly $20.40. All the code ran inside an isolated sandbox with no access to outside systems.
How much do the safeguards actually block
When asked directly for something malicious, GLM-5.3 refuses about 95% of the time, and its simulated attack success rate is 0. That changes once you try the following approaches:
| Method | Bypass success rate |
|---|---|
| Invent a seemingly legitimate pretext | 64% |
| Pre-fill a chain of reasoning | 92% |
| Use "ablation" to strip out the refusal mechanism | 100% |
Ablation requires modifying the model's weights, something only possible with an open-weight model. Anthropic says its team had never done this before; the process took roughly 2,200 GPU-hours and cost about $4,400, and it dropped the refusal rate from 95% down to somewhere between 3% and 14%. As a point of comparison, the guarded Claude models were not broken in the same tests.
The report makes three recommendations: give defenders access to frontier models that are equally or more capable; have governments run safety testing on high-capability models; and have open-weight model developers add corresponding safeguards. This outlet found no public response from Zhipu to the report.
A different assessment puts the numbers differently
CAISI (the Center for AI Standards and Innovation), under the U.S. Department of Commerce, also published its own evaluation of GLM-5.3's cyber capability on September 17. Its conclusion: GLM-5.3 is currently the most capable open-weight model for cyber tasks, but it clearly lags U.S. frontier models, by roughly four months overall.
The two reports' numbers diverge sharply. On CAISI's version of ExploitBench, GLM-5.3 scored 61.1%, the best U.S. frontier model scored 100%, and the previously best Chinese model scored 32.2%; on SEC-Bench Pro, it was 40.4% versus 90.2%. Anthropic's 12% figure comes from its own test configuration, so the two sets of results aren't directly comparable.
The two reports agree on the ranking: GLM-5.3 leads the open-weight field but still trails the strongest closed U.S. models. Where they differ is how to read that gap. CAISI emphasizes that GLM-5.3 is "still about four months behind"; Anthropic emphasizes that this level of capability can already be downloaded and modified by anyone, and that safety layers can't meaningfully constrain it once that happens.
For developers and companies in China, the report's most direct impact may land on overseas distribution channels. GLM-5.3 has already been fairly widely listed on overseas API distribution and hosting platforms; if more U.S. institutions adopt Anthropic's framing, overseas cloud platforms and enterprise customers may tighten their review of Chinese open-weight models before listing them. For now this remains speculation — no platform has taken action yet.
Sources: Anthropic research report, CocoLoop, the U.S. NIST-affiliated CAISI assessment announcement; ExploitBench success counts, bypass rates and ablation costs follow Anthropic's report methodology, while the CAISI figures follow its own test configuration.