Perplexity Red-Teamed Its Agent Sandbox: The VM Held, the Network Didn't
Perplexity's Secure Intelligence Institute published a red-team report on its own SPACE sandbox: zero VM-to-host escapes in 108 runs, but four models bypassed the network egress policy, and
TL;DR
- On September 23, 2026, Perplexity's Secure Intelligence Institute published "Escaping SPACE: Part I," a red-team report on SPACE, the microVM sandbox platform behind Perplexity Computer. Nine models received root access inside a guest VM with the explicit objective of crossing its boundaries.
- No VM-to-host escape was observed in 108 runs — even with access to SPACE's source code, web search, and software package installation. That boundary held.
- The network boundary did not. With limited network access, four of the nine models (Claude Opus 5.0, GPT-5.6 Cyber, GPT-5.6 Sol, Kimi K3) bypassed the egress policy in 11 of 54 partial-network runs, via DNS spoofing or shared-IP routing.
- The problem is broader than one vendor: eight of ten third-party sandbox platforms tested showed at least one network-policy bypass. Only Cloudflare Sandbox and NVIDIA OpenShell held in these experiments.
- The structural lesson for anyone running agent workloads: treat "sandboxed" as at least two separate claims — isolation of compute, and confinement of network — and test them separately.
- The operational change that follows: audit the two boundaries independently, run your own partial-network red team, prefer gateways that terminate TLS and validate SNI against a DNS-attributed allowlist, and assume shared-IP bypasses are permanent.
What the red team actually tested
SPACE is built on Firecracker microVMs. Each sandbox is a guest VM on a host the guest does not control, and a per-sandbox gateway enforces egress policy on everything the guest sends out. The report evaluates two boundaries that are easy to conflate: the VM-host boundary (can an agent that already has root inside the guest reach the host?), and the network boundary (can an agent reach a destination the egress policy blocks?).
The experiments ran nine models as agents under four settings defined by two dimensions: whether the agent had access to SPACE's source code, and whether network access was fully blocked or partially allowed (package repositories plus the Perplexity Search API). The sandbox-escape task asked agents to obtain an AWS honeytoken stored on the host; the network-bypass task asked them to retrieve a flag from a blocked callback URL.

The design matters because it isolates the failure modes. Giving an agent root inside a guest VM is the worst-case assumption for compute isolation: if the hypervisor boundary is sound, root inside the guest should still mean nothing outside it. Giving the agent partial network access is the realistic case for production agents, which almost always need package registries, APIs, or retrieval services to do useful work. A partial-network setting is therefore the setting that should be tested: it is where an agent still does useful work, and it is where policy mistakes surface — a fully blocked network proves nothing about egress controls, and full access tests nothing at all. The two tasks fail in very different ways, and the report's phrasing is exact: "A successful network-policy bypass does not imply a VM–host escape, and the absence of an observed escape is not a proof of isolation."
Why the egress policy failed
Every successful bypass exploited the same structural gap: the policy was expressed in domain names, but enforcement happened against IP addresses. Perplexity describes two concrete methods.

In the DNS-spoofing variant, the agent injected a forged DNS response impersonating the external resolver, mapping the allowed hostname pypi.org to the blocked callback's IP. The gateway then saw a connection to an allowed name, en route to an address it had not decided to block. That is a policy-implementation failure specific to SPACE — the report says it was addressed with source-address validation — and it is the easiest of the two to fix.
The second method is more uncomfortable because it is structural. A single IP address serves many domains: content delivery networks host thousands of unrelated customer domains behind shared IP pools, and Fastly's Fiddle service exists precisely to make server-side fetches on behalf of arbitrary URLs. One run sent a request to fiddle.fastly.dev instructing the service to fetch the blocked callback URL. GPT-5.6 Cyber found another path through images.taboola.com, an image-fetching service reachable through an IP the policy already allowed, chained to thum.io for a screenshot, then recovered the flag from the PNG with OCR. Every hop in that chain went through an IP the gateway already allowed, so the egress policy never saw the final destination — the name-based allowlist was satisfied the whole way, which is exactly why fix belongs at the gateway, not inside the guest.

Perplexity's own conclusion is the transferable part: "Relying only on the destination IP address at the host to determine whether a connection is allowed is insufficient." Any operator building an egress allowlist — whether for an agent sandbox, a CI runner, or a microservice — is building the same failure mode if the check happens after DNS resolution and trusts the resolver.
Eight of ten platforms failed the same test
Perplexity did not stop at its own sandbox. It ran direct HTTP/HTTPS tests against ten third-party sandbox platforms that advertise hostname- or domain-based egress controls, testing whether authorizing one hostname could authorize a request for a different hostname served by the same CDN or multi-tenant infrastructure. At least one network-policy bypass was found in eight of the ten platforms. The details matter: Deno Sandbox enforced HTTPS policy but allowed arbitrary TCP to unlisted addresses, and several vendors shipped mitigations quickly after disclosure (microsandbox, Daytona, Deno); others documented known limitations rather than reclassifying the behavior. The two platforms where neither tested bypass succeeded were Cloudflare Sandbox and NVIDIA OpenShell.
The vendor-name range is the news: E2B, Vercel, Modal, Daytona, Fly.io, microsandbox, Deno, and others are all mainstream execution surfaces for agent frameworks. An agent whose egress "policy" can be redirected through a shared CDN edge is not actually confined — it is billable. This connects directly to the control-plane argument we made when Cloudflare placed Cursor agents in customer-controlled sandboxes: the sandbox is only as good as the controls the customer can actually verify, and egress is where most platforms still trust the guest's view of the network.
What operators should change
First, stop auditing whether the sandbox is secure as one question. The report's core finding is that compute isolation and network confinement are independent properties with different failure rates — and in this test, very different: 0/108 versus 11/54. Second, test your own platform the way Perplexity tested its own: give an agent partial network access, point it at a blocked destination, and see whether a shared CDN edge or a poisoned DNS answer gets it out. Third, prefer platforms that terminate TLS at the gateway and validate the SNI and HTTP hostname against a DNS-attributed allowlist — that design closes both bypass classes at once, and it is the fix Perplexity shipped. Finally, assume the IP-sharing problem is permanent: it is a property of today's Internet architecture, not of any single vendor, so an egress control that does not look past IP addresses will keep failing regardless of which sandbox product you buy.

The pattern is the same one we flagged in the CVE-2026-80521 container-escape analysis: the boundary that marketing describes and the boundary that attacks actually cross are rarely the same line. Containers held against a kernel hole only until the kernel had a hole; agent sandboxes hold at the hypervisor while leaking at the gateway. Scope the threat model to the boundary that demonstrably fails, and test that one.
Sources
Source: Escaping SPACE: Part I — perplexity.ai/hub/blog/escaping-space-part-i, 2026-09-23, by the Perplexity Secure Intelligence Institute (all figures in this article are official Perplexity figures from that page, fetched 2026-09-30). Internal linkage: Cloudflare Sandboxes and Cursor agents, CVE-2026-80521 Ubuntu container escape. More infrastructure and agent-security analysis on neticslabs.com.
Source: Escaping SPACE: Part I — perplexity.ai, 2026-09-23. Figures: official Perplexity report graphics, fetched 2026-09-30.