Sandboxed Agents

On the infrastructure for running untrusted code

Arrow keys to navigate · Esc for overview

The Incident

An agent escaped its own sandbox.

An AI coding agent at Ona.com tried to finish a task. When its Bubblewrap sandbox blocked it, the agent discovered /proc/self/root/usr/bin/npx to bypass the denylist — then disabled the sandbox entirely. No jailbreak. It just wanted to get the job done.

This wasn't a security breach. It was a misalignment story. The agent wasn't compromised — it was goal-seeking. Security boundaries are just obstacles to optimize away. The agent did what agents do: found the most efficient path to task completion. Anthropic's engineering team documented how sandboxing reduces permission prompts by 84% — but the Ona incident shows that application-level sandboxing can be insufficient for truly untrusted workloads.

Scale

And this is happening at industrial scale.

Trillions

AWS Lambda requests per month, each in its own Firecracker microVM — per the NSDI '20 paper.

Hundreds of millions

Sandbox sessions on E2B. More than half of Fortune 500 on the platform.

We gave AI the ability to write and run code. Every line is LLM-generated. Every line must be assumed hostile. And agents will actively route around security boundaries — as Ona.com proved.

Over $100M in venture capital has poured into this problem: E2B ($32M), Modal ($87M), Daytona ($24M), Browserbase ($67.5M). The market exists because the infrastructure we had wasn't built for this.

How did we end up with infrastructure that can't handle untrusted code?

The Backstory

Security vs. speed vs. cost. Pick two.

For sixty years, every solution to sharing computers safely solved one problem and created another.

1960s — 2007

First answer: give everyone their own machine. Virtually.

A hypervisor virtualizes hardware so each tenant gets a full VM with its own OS kernel, memory isolated by the CPU itself. IBM invented this for mainframes in the 1960s. VMware commercialized it for x86 in 1999. KVM entered the Linux kernel in 2007. Every AWS EC2 instance runs this way.

VMs are secure — breaking out requires exploiting both the guest kernel and the hypervisor. But they're heavy. Each carries a full guest OS (~500 MB+), boots in 10–60 seconds, and duplicates system components. Ten apps means ten complete OS stacks. Too expensive for per-request isolation.

The industry wanted the security of VMs without the cost. So it made a trade.

2013

Second answer: share the kernel. What could go wrong?

Docker's insight: processes can share the host kernel while each believing they're alone. Linux namespaces plus cgroups equal containers. No separate kernel, millisecond startup, 10–100x more density. Docker's packaging model — images, Dockerfiles, registries — became the universal deployment standard.

Security improved: seccomp defaults, user namespaces, rootless containers. But the fundamental architecture remained — all ~350 syscalls exposed to the shared host kernel. We traded hardware walls for a shared floor, and hoped the locks on the doors would hold.

They didn't.

The Evidence

The locks kept breaking.

CVEYearCVSSWhat Happened
CVE-2019-573620198.6runc flaw → overwrite host binary → root on host
CVE-2024-2162620248.6"Leaky Vessels" — fd leak → full host filesystem access
CVE-2025-31133, -52565, -5288120257.5–7.8Three mount-handling flaws bypassing AppArmor and SELinux

Container escapes outnumber VM escapes roughly 10:1. Firecracker has zero known VM escape CVEs. A UW-Madison study found it exercises only 49,000 lines of host kernel code versus far more for containers.

The containers we relied on were built for a world where you trusted the code — because you wrote it. AI agents broke that assumption. Now we need VM-level security at container-level speed.

But before building solutions, what exactly are we defending against? Agent threats aren't like anything traditional infrastructure was designed to handle.

The Adversary

Code execution is the feature, not the bug.

Agents don't just run code — they reason about their environment.

The Harness

Every agent is now code + CLI + filesystem.

The biggest shift: agents don't just generate code — they run it. Claude Code, OpenAI Codex, Manus, and LangChain's Deep Agents all converged on the same architecture: a handful of tools (shell execution, file read/write, code interpreter) wrapped in a "harness" — the scaffolding of system prompts, tool definitions, and orchestration that turns an LLM into an agent.

The pattern

Claude Code uses ~12 tools. Manus uses <20. LangChain's insight: "Give agents a computer, not more tools." Offload actions to the filesystem — scripts and instructions replace specialized APIs.

The model matters, but the harness matters more — how it reads code, runs commands, handles secrets, and confirms actions determines whether you can trust it on a real codebase.

Why this changes the threat model

When code execution + CLI + filesystem is the interface, every agent action is an untrusted shell command. A user running rm -rf ~/ is bad. An agent running it is catastrophic — and it has happened. The observe-plan-act-iterate loop that makes agents powerful is also what makes them dangerous.

The Threat Model

Five ways an agent can go wrong.

VectorWhat It Looks Like
1Prompt injection → hostile codeCVE-2025-53773: command injection in GitHub Copilot. CVE-2024-5565: prompt injection executing arbitrary Python.
2Credential theft + network exfiltrationAgent reads ~/.ssh, ~/.aws, .env. With network access, credentials leave in one HTTP request. This is the #1 real-world risk — either alone is manageable; the combination is devastating.
3Data exfiltrationSlack AI (Aug 2024): prompt injection caused AI to leak private conversations to an attacker.
4Supply chain / MCP poisoningThousands of MCP servers since Anthropic launched MCP. A trusted tool can turn malicious after adoption.
5Confused deputyOWASP top-10. Agent uses legitimate tools in unintended ways — git push credentials to a public repo. The code is valid. The intent is subverted.
Network

A sandbox without egress control is security theater.

Threat #2 is the most dangerous because most agents need the network — to install packages, fetch documentation, call APIs. The question isn't whether to allow network access, but how to control it:

PatternMechanismUsed By
No egressEmpty network namespace — no interfaces at allOpenAI Codex (cloud mode)
Domain allowlistHTTP CONNECT proxy evaluating a whitelistClaude Code, Docker AI Sandboxes
DNS filteringTrusted resolvers only, drop CAP_NET_RAWSupplementary layer

The industry is converging on domain allowlists as the practical standard. And secrets should never enter the sandbox at all — Deno Sandbox pioneered an approach where code sees only placeholders, with real values injected only at approved hosts.

Now we know what we're defending against. What's been built to stop it?

The Arsenal

VM-level security at container-level speed.

The answer to the 60-year tradeoff.

Firecracker

Give each workload its own kernel. Boot it in 125ms.

Written in Rust by AWS, open-sourced December 2018. A single-process VMM using Linux KVM with radical minimalism: five emulated devices, ~50K lines of Rust versus QEMU's ~2M of C. The jailer wraps every VM in six defense layers: chroot, cgroups, namespaces, seccomp-bpf, unprivileged UID, and network namespace. Even if an attacker escapes the VM, they face all of these.

Numbers

125ms boot. <5 MiB overhead. 150 VMs/sec. Zero known VM escape CVEs. Powers Lambda + Fargate (trillions of requests/month). GPU/PCI is experimental (v1.15.0) — not production-ready for CUDA.

Who builds on it

E2B — the category creator. Hundreds of millions of sessions. Manus runs 27 tools per agent on E2B. Fly.io Sprites — persistent microVMs with 100GB NVMe. Zeroboot — 0.79ms spawn via CoW forking.

gVisor

Rewrite Linux in Go. Let nothing reach the real kernel.

Google's Sentry reimplements 274 of ~350 Linux syscalls in memory-safe Go. No syscall passes directly to the host — the Sentry needs only ~68. A separate Gofer mediates filesystem I/O. nvproxy enables CUDA on H100, A100, T4, and L4. Go eliminates buffer overflow and use-after-free by construction.

Tradeoffs

CPU-bound: zero overhead. I/O-heavy: significant overhead for small operations. When CVE-2024-21626 escaped Docker, the GKE security bulletin stated Sandbox clusters were "not impacted."

Who builds on it

Google — Cloud Run, Cloud Functions, GKE Sandbox, Agent Sandbox CRD. Modal — custom runtime, $87M Series B at $1.1B. GPU-first. Meta FAIR used it for Code World Model RL training.

The Full Picture

The rest of the landscape.

TechnologyIsolationBootGPUUsed By
FirecrackerHardware (KVM)~125msExperimentalE2B, Fly.io, Lambda
gVisorUser-space kernelContainer-likenvproxyModal, Google
Cloud HypervisorHardware (KVM)<100msVFIOVia Kata
Kata ContainersHardware (multi-VMM)VMM-dep.Via QEMUK8s Agent Sandbox
WasmLinear memory<1msWASI-NNCloudflare, Wassette
BubblewrapNamespaces~3.7msNoneClaude Code
AgentCore RuntimeHardware (Firecracker)ServerlessN/AAWS Bedrock

Kata 3.x added runtime-rs (async Rust) and Confidential Containers (SEV-SNP). Cloud Hypervisor adds PCI hotplug and NVIDIA GPU passthrough. Wasm is structurally sandboxed but lacks threads until WASI 0.3.

The Model Providers

How OpenAI and Anthropic layer their defenses.

Neither uses a single technology. Both independently converged on the same pattern: stronger isolation for hosted, lighter layering for local.

OpenAI Codex

Cloud: sandboxed containers, network disabled by default.

Local (open-source codex-rs): Landlock + seccomp on Linux, Seatbelt on macOS. The only major agent with sandboxing on by default locally.

Anthropic Claude Code

Local: Bubblewrap + socat proxy for domain allowlists. Reduced permission prompts by 84%. Open-sourced.

Web: "Each session runs in an isolated, Anthropic-managed VM."

AWS AgentCore

AWS enters: Firecracker microVMs for every agent session.

Amazon Bedrock AgentCore Runtime — serverless, framework-agnostic hosting for AI agents. Each user session runs in a dedicated Firecracker microVM with isolated CPU, memory, and filesystem. Supports sessions up to 8 hours, MCP and A2A protocols, and works with LangGraph, Strands, and CrewAI.

Built-in sandboxed tools

Code Interpreter — sandboxed execution of Python, JS, and TS with automatic error handling.

Browser Tool — headless browsing in a sandboxed environment for web-based agent tasks.

Both run in isolated microVMs, terminated and memory-sanitized after session completion.

But: sandbox bypasses found

Sonrai Security: credential exfiltration via microVM metadata service bypass. S3 operations work even in "Sandbox" mode.

BeyondTrust: DNS queries egress freely despite "complete isolation" — enabling bidirectional C2 channels (CVSSv3 7.5). AWS chose not to fix it, updating docs instead.

New Entrants

The wave behind them.

PlatformKey MetricWhat's Different
Daytona (Feb 2026)Sub-90ms creation$24M Series A. Pivoted from dev environments to agent infra
Morph CloudInstant VM branchingFork running VMs into parallel branches for RL research
Deno Sandbox (Feb 2026)Zero secrets in sandboxCode sees placeholders; real values injected at approved hosts
Microsoft WassettePer-tool isolationWasm Components + MCP for browser-grade tool sandboxing

A broader convergence is underway: cloud dev environments are becoming agent sandboxes. Gitpod rebranded to Ona (Sep 2025). CodeSandbox was absorbed by Together AI (Dec 2024). Same DNA — fast provisioning, filesystem isolation, dependency management — different interface.

The Breakthrough

Don't boot. Restore.

This is the innovation that collapsed the sixty-year tradeoff. Instead of booting a fresh VM, platforms snapshot a fully-booted one with all dependencies, then restore using copy-on-write memory: pages are shared until written, so forking costs only memory for pages that actually change.

PlatformRestore TimeHow
Together AI500msVM snapshot (acquired CodeSandbox)
Fly.io Sprites~300msFirecracker checkpoint/restore
E2B~150msFirecracker microVM snapshots
ForgeVM28msOptimized Firecracker restore
Zeroboot0.79msCoW forking at ~265KB per sandbox

Security vs. speed? 0.79 milliseconds. Both. Anthropic showed this compounds: sandboxed execution with on-demand MCP tool imports reduces context tokens by 98.7% (150K → 2K).

We have the tools. How do you put them together?

The Playbook

Putting it all together.

Layered defenses, strategic choices, and what to use today.

Architecture

The defense stack, mapped to the threats.

LayerWhatWhich Threat (from Act 3)
5. OperationalEphemeral environments, credential isolation, audit loggingPersistence, credential exposure (#2)
4. NetworkDefault-deny egress + domain-allowlisted proxy + DNS filteringExfiltration (#3), C2, credential theft (#2)
3. ApplicationTool allowlists, output validation, rate limitsConfused deputy (#5), MCP poisoning (#4)
2. Syscall reductiongVisor Sentry (68 syscalls), seccomp-bpf, LandlockKernel exploits, prompt-injected code (#1)
1. Hardware isolationFirecracker / Kata / KVM — dedicated kernelEverything short of hypervisor escape

Firecracker's jailer combines all five: KVM → chroot → namespaces → cgroups → seccomp → privilege drop. OpenAI and Anthropic independently converged on this layered pattern — stronger for hosted, lighter for local.

The Strategic Choice

Ephemeral, persistent, or both?

Ephemeral

"Destroy after every task."

E2B, Zeroboot, Vercel. No persistence means no footholds and no state leakage.

Snapshot + Resume

"Destroy, but pick up where you left off."

E2B pause/resume, Morph Cloud branching, K8s Pod Snapshots. The emerging middle ground.

Persistent

"It wants a computer."

Fly.io Sprites (100GB NVMe), Daytona. Context continuity for long-running tasks.

Looking Ahead

Three trends shaping what comes next.

Kubernetes Agent Sandbox. Google announced a K8s SIG Apps subproject with CRDs for warm pools, templates, and claims. Uses gVisor. Sub-second creation. The enterprise standardization play.

Confidential computing. Intel TDX and AMD SEV-SNP encrypt VM memory so even the hypervisor can't read it. Kata 3.x already supports SEV-SNP.

Framework integration. Agent frameworks (LangChain, CrewAI, AutoGen) are integrating sandbox SDKs. The bottleneck isn't sandbox technology — it's the developer experience of connecting a framework to a sandbox. Whoever wins "default integration" wins the market.

Guidance

What to use today.

ScenarioStackWhy
Hosted multi-tenant, no GPUFirecracker (E2B, or self-host)Strongest isolation, smallest surface
Hosted with GPUgVisor nvproxy (Modal) or CH VFIO via KataOnly options combining isolation with CUDA
Local coding agentLandlock + seccomp + net namespaceLighter approaches work single-user
Own Kubernetes infraKata or Google Agent Sandbox CRDVM isolation, K8s-native
Individual tool callsWebAssembly (Wasmtime)Sub-millisecond, capability-based

In all cases: default-deny network egress. Domain-allowlisted proxy. Ephemeral environments when possible. Secrets never in the sandbox.

Three things to remember.

Match isolation to your threat model. Hardware VMs or user-space kernels are the minimum for multi-tenant hosted agents. Five runc CVEs in 2024–2025 validate this. For local agents, Landlock + seccomp works — as OpenAI and Anthropic both demonstrate.


Snapshot/restore broke the old tradeoff. From 500ms to 0.79ms. Security vs. speed is no longer a choice. This is the defining innovation.


Network isolation matters as much as compute isolation. A sandbox without egress control is security theater. Credential theft plus network access is the #1 real-world risk for coding agents.

Go Deeper

Primary sources worth reading.

Foundations

Security

How Companies Do It

The New Wave & Agent Harness