Claude Geolocates Photos Within 37km, Beating Humans

Anthropic's Frontier Red Team and threat intelligence group published an evaluation on September 10 measuring how far models can go on military and intelligence tasks. The tasks split into two categories: intelligence work, including linking accounts across social platforms to the same person and inferring locations from photos and text; and conventional weapons development, including drone terminal guidance, payload delivery, and autonomous navigation when GPS is jammed.

The models tested were Anthropic's own Mythos Preview, Mythos 5, Opus 5, and Sonnet 5, plus the open-weight Kimi K3. The report's overall conclusion:

"For some tasks in military and intelligence domains, models could do things that, historically, only a set of scarce, highly-trained human experts could do."

Finding places: photos and text both work

Photo geolocation used 6,000 outdoor photos from a public Flickr dataset, stripped of metadata, with reverse image search disallowed. Mythos Preview's median error was 37.0 kilometers, with 23.7% of photos located to within 1 kilometer. The comparison group — competitive players of the geography-guessing game GeoGuessr — had a median error of 151 kilometers. Opus 5 came in at 181 kilometers, while Sonnet 5 and Kimi K3 both landed around 385 kilometers.

Text-based geolocation used a corpus of tweets covering 1,697 users, with models given access to search tools. Mythos Preview, Mythos 5, and Opus 5 all had median errors between 20 and 22 kilometers, and 135 users (about 8%) were located to within 1 kilometer by at least one model.

The account-linking task spanned platforms including WhatsApp, Instagram, and Facebook, with 200 questions total. The report's efficiency comparison: a 37,000-word document took a model about 11 minutes to analyze, versus 2.5 hours for a human just to read through it once.

Hitting targets: easy when still, chaotic when GPS is spoofed

Drone terminal guidance covered 9 scenarios, including moving targets, camouflaged targets, and roadside obstacles:

ModelStationary high-contrast targetMoving targetAll 9 scenarios combined
Opus 580%47%20%
Mythos Preview70%20%13%
Mythos 553%17%10%
Kimi K315%1.6%1.6%
Sonnet 55%0%0.7%

Payload delivery counted a hit as landing within 5 meters of the target. On stationary targets, Opus 5 and Mythos 5 came close to a perfect record, with median errors of 20 to 30 centimeters. When the target moved at 3 meters per second, Mythos Preview and Opus 5 scored 77% and 76% respectively; add wind, and only Opus 5 kept any meaningful hit rate, at 28%.

GPS navigation was tested under three conditions. With a normal signal, every model except Sonnet 5 could reach the destination. When GPS was jammed, Opus 5 stopped 15 to 20 meters from the target, with 33% of runs landing within 5 meters, while the Mythos series averaged around 30 meters off. Add stealthy spoofing and continuous drift, and every model lost accuracy.

Kimi K3 as the open-weight benchmark

The report describes open-weight models as generally trailing the frontier, with scores typically falling between the Sonnet and Mythos tiers, but often already carrying concerning capabilities. It adds that frontier models themselves can clear this bar easily, and the open-weight ecosystem is catching up fast — "this gap should not be mistaken for safety."

On policy, the report argues for preserving a chip advantage for "democratic nations" to limit the pace of rivals' AI progress. The same week, Anthropic also published a threat intelligence report naming several Chinese labs it says are distilling Claude at scale.

These are floor scores, not ceilings

The report lists several caveats: the data is synthetic or simulated, the drone tests ran on simplified visual rendering rather than real hardware, and the models operated in an isolated sandbox with no internet access and no full off-the-shelf codebases. Anthropic frames these results as a "floor, not a ceiling" on capability.

On the defensive side, Anthropic says that after finding Claude being misused for weapons development, it has deployed a new classifier to detect and block related requests. The report doesn't disclose the classifier's block rate or false-positive rate.

Sources: Anthropic Frontier Red Team report, CocoLoop, Anthropic threat intelligence team public materials; hit rates and geolocation errors follow the report's stated simulated-test methodology.