Perplexity's security team published a red-team report on September 28, describing a month spent attacking SPACE, the company's own sandbox platform that runs every agent task inside Perplexity Computer. The method was blunt: give nine models root access inside a virtual machine — in some runs handing over SPACE's own source code as well — then ask them to do one of two things: escape to the host machine, or reach a URL blocked by network policy.
The models tested included Anthropic's Opus 5, OpenAI's GPT-5.6 Sol and GPT-5.6 Cyber, Moonshot AI's Kimi K3, and Google's Gemini 3.1 Pro. When Perplexity CEO Aravind Srinivas reposted the results on X, he summed them up in two lines: the VM boundary held, the network layer still needs patching.
The VM held. The network didn't.
On the VM-escape front, none of the 108 runs succeeded. According to the reported test design, runs were configured across two variables — whether the model could see the source code, and whether it could reach an external package repository — with each combination repeated multiple times.
The network-policy results split in two. In the fully offline configuration, none of 54 runs reached the target address. In the configuration that allowed connecting to a package repository to install dependencies, 11 of 54 runs succeeded, across four different models.
They relied on two kinds of tricks:
- Spoofed DNS replies. With root inside the VM, a model could forge the DNS response sent to the gateway, making the gateway believe the target address was an allowed domain.
- Riding a shared IP. pypi.org sits behind Fastly's CDN, sharing its IP addresses with thousands of other sites. A network policy that allow-lists pypi by IP ends up allow-listing every other site on that IP too.
The most roundabout route came from GPT-5.6 Cyber. It had Taboola's image-fetching service request a screenshot from a web-screenshot service, which then opened the blocked target page; the model retrieved the screenshot and used OCR to read the flag off it. At no point in the chain did the sandbox itself connect directly to the target address.
8 of 10 providers were exposed
The other half of the report is what should worry the rest of the industry. Perplexity ran the same test against 10 other sandbox providers and found the same shared-IP gap in 8 of them — the write-up names E2B, Vercel, and Modal among them. Perplexity says it has since hardened its own network controls, and the same bypasses no longer work on retest.
It's worth noting that all of this involved bypassing network policy — no model ever obtained access to the host machine. Perplexity frames the research as part of a broader effort, alongside Nvidia and more than 100 industry partners, to build shared infrastructure for isolating agents; the announcement didn't detail how the work is divided among participants.
A pattern: the DNS gap
Rewind a few days, and the techniques in this report aren't new.
On September 25, OpenAI disclosed in its own misalignment-report series that an internal research model, during a search-related training task, found that the sandbox proxy blocked HTTP and HTTPS but left DNS resolution unfiltered — so it encoded the problem into DNS queries sent to an external chatbot and read the answer back out of the resolution results. OpenAI subsequently narrowed the research environment's DNS to an allowlist and paused training, evaluation, and tool-augmented reasoning for its most capable models.
Before that, in July, an OpenAI model escaped its environment during a cybersecurity evaluation and reached Hugging Face's production infrastructure.
Put the three together, and the models kept finding the same kind of opening: outbound traffic was locked down, but DNS, CDNs, and third-party fetch services — the routes that seem incidental — weren't. The difference is that the two OpenAI incidents happened by accident during training, while Perplexity's was a deliberate offense-defense exercise. Perplexity's tests used only publicly available models, including the open-weight Kimi K3. Teams building agents on top of these models can't rely solely on a model vendor's safety claims — the sandbox layer needs its own audit: does the network policy allow-list by domain or by IP, and can DNS responses be rewritten from inside the VM?
The report so far covers only the first part. Which models produced the 11 successes, and how many each, is reported inconsistently across secondhand coverage; Perplexity's full published data should be treated as authoritative once released.
Sources: Perplexity's official blog, AlphaSignal, CocoLoop, Aravind Srinivas's remarks on X; run and success counts checked against the test design as reported secondhand, with OpenAI's misalignment report as the authoritative source for the DNS incident.