Anthropic published a threat intelligence report on September 10 that sorts Claude misuse cases identified between December 2025 and August 2026 into seven categories: cyber operations, influence manipulation, surveillance, conventional weapons development, biological misuse, fraud and scams, and illicit distillation.
The conventional-weapons category holds six cases in total. By the attribution the report gives, three trace back to China, two to Russia, and one to Yemen. The Yemen case is the one described in the most detail.
One unit, three projects
The report describes a weapons-development unit in northern Yemen running three projects at once: a guided rocket, a multi-stage ballistic missile with a range beyond 2,000 kilometers, and a family of hypersonic glide variants.
What the unit used Claude Code for was guidance, navigation and control software — the GNC layer, in industry shorthand. The report lists four specific tasks: writing control and position-estimation code, tuning parameters, building firmware, and running flight simulations. Anthropic's own wording is that the people involved used Claude Code to handle work that would otherwise have required a software engineer.
Two techniques were used to get past the safety mechanisms. One was disguising the software's military purpose, framing requests as generic control-algorithm problems. The other was splitting the whole workflow across separate, unconnected sessions, so that each session only ever saw one piece and none could reconstruct the full plan.
What exposed the operation was a failure. According to the report, a test launch of the guided rocket did not succeed, and when the operator went back to ask Claude what had gone wrong, that query tripped a safety alert.
Anthropic says its safety mechanisms blocked many of the requests involved, though not all of them, and that the accounts involved have been permanently banned.
One point needs clarifying: the report does not name which side the unit belongs to. "Northern Yemen" is the geographic description the report gives; the identification with the Houthi movement comes from media inference, and Anthropic has not confirmed it. The report also states there is no evidence any of the three projects produced a usable weapon.
Surveillance is the category with the bigger numbers
Of the seven categories, surveillance is the one most likely to get lost in scale. The report cites a platform in Mali called Lakana 360 that covers roughly 25 million SIM cards across three nationwide mobile carriers. That scale goes beyond an individual overstepping the rules — it describes systemic infrastructure.
The illicit-distillation category centers on several Chinese labs. The Alibaba-linked figure cited is 151 million interactions over three months, and the report describes this category as "industrial-scale covert operations." That part has already been reported separately elsewhere.
The report also lays out a detail about which models were involved: most of the misuse involved older models such as Claude Opus 4 and Sonnet 4.5. According to the report, the biological safeguards on those older models were "designed mainly to prevent novices from accessing content about known bioweapons," while the newer Fable- and Mythos-tier models carry tighter restrictions, narrowing even broad, dual-use biological-research queries.
Two documents, same day
The threat intelligence report wasn't the only thing Anthropic released on September 10. The same day, its frontier red-teaming and threat intelligence teams published a capability evaluation testing how far models can go on military and intelligence tasks — including photo geolocation, account linking, terminal-phase drone guidance, and autonomous navigation when GPS is jammed. Its conclusion: on some tasks, models can now do things that used to require a small number of trained specialists.
Read together, the two documents run in reverse chronological order. The capability evaluation answers "can it be done"; the threat intelligence report answers "is someone already doing it." The latter's coverage window starts in December 2025, more than half a year before the former's testing period.
Looking at Anthropic's public cadence over the past year, this line of disclosure has been intensifying — from early, sporadic notices of account bans, to single-incident disclosures like the Mexican water-system case, to this latest year-in-review organized by category and attribution. Whether this level of disclosure becomes standard practice across the industry is still an open question; none of Anthropic's peers has published comparable material so far.
Sources: Anthropic threat intelligence report, CocoLoop, UDN, unwire.hk.