Firecracker vs gVisor vs Kata Containers for Agent Sandboxing
Three isolation layers solve different parts of the agent sandbox problem.

An AI agent asked to fix a bug reads the codebase, runs pip install on whatever package the dependency tree pulls in, executes shell commands, and makes outbound calls to whatever host its tool-use decides to hit, all before a human looks at a single diff. That sequence is the whole problem. A web server or a CI job runs code a team wrote and reviewed; an agent runs code nobody has read yet, pulled from a public registry at the moment of execution, with the shell as an open microphone.
The threats that follow from this are not exotic, and treating them as edge cases is where a lot of sandboxing designs go wrong. Prompt injection, typosquatted packages with malicious post-install hooks, and supply-chain attacks against npm and PyPI are the expected inputs to an agent sandbox, not the rare failure mode. A typosquatted package with a post-install script that scans the local network isn't a theoretical adversary model. It appears in production traffic.
That changes the question an operator has to ask. Container security, in its traditional form, asks how to harden the box: tighter seccomp profiles, read-only file systems, a leaner base image. Agent sandboxing has to ask a blunter question: once a sandbox is fully compromised, root inside it, hostile code running with whatever privileges the process has, what does the attacker reach next? Is a neighboring tenant's workload one step away, or does the attacker need a second, independent exploit to get there? Blast radius, defined that precisely, is the only metric that matters once you accept that compromise of the sandbox itself is a routine, expected event rather than a rare failure. That's the lens the rest of this comparison uses, and it's why the three technologies people keep putting side by side don't actually answer the same question.
The layer confusion that makes most comparisons of these three technologies useless
Most write-ups line up Firecracker, gVisor, and Kata Containers in a feature table as if they were three competing answers to the same question. They aren't, and a comparison built on that premise can't settle anything, because it's comparing things that don't sit at the same layer of the stack.
Firecracker is a hypervisor: a userspace virtual machine monitor that uses KVM to boot a virtual machine with its own guest kernel. Cloud Hypervisor and QEMU live at that same layer, doing the same fundamental job with different tradeoffs. gVisor is a different kind of thing entirely. It is neither a hypervisor nor a traditional container runtime; it's a userspace kernel that intercepts syscalls through a component called the Sentry, inserting itself between the container and the host kernel rather than booting a separate one. Kata Containers is an OCI runtime (the component Kubernetes or containerd calls to create a container) that answers that call by booting a VM through one of several pluggable hypervisors: QEMU, Cloud Hypervisor, or Firecracker. Choosing Kata doesn't remove the hypervisor decision. It just moves it one layer down.
Once that's clear, the real decision is two sequential questions rather than one three-way pick. First: does the untrusted workload get its own kernel, or does it share the host's? Second, if it gets its own: which hypervisor boots it? The actual decision splits into two sequential questions: does the untrusted workload get its own kernel or share the host's, and if it gets its own, which hypervisor boots it.
The evidence backs this up in a way that's easy to miss until you go looking for it. When one comparison fingerprinted the isolation stacks actually running behind six production sandbox vendors, what showed up was Docker-based containers, gVisor, Firecracker microVMs, and Cloud Hypervisor.
What the kernel-sharing boundary means for isolation strength
Strip away the branding and the isolation question comes down to one thing: does untrusted code share a kernel with anything the operator cares about? The kernel is the single enforcement point for every resource boundary on a Linux host, memory, processes, file descriptors, the works, so whatever shares it shares the fate of everything running under it.
A container narrows what a process can touch using namespaces and seccomp filters, while a VM uses a hypervisor like Firecracker (a userspace virtual machine monitor that uses KVM to boot a VM with its own guest kernel) to isolate a guest. They aren't, and the difference is in what actually enforces the boundary rather than in what the boundary claims to do. Namespaces and seccomp are software checks running inside the same kernel the attacker is trying to break, which means their guarantee is only as good as that kernel's freedom from exploitable bugs. A hypervisor boundary is enforced by hardware virtualization instead, so an attacker who fully compromises the guest still has to escape the guest kernel and then separately escape the hypervisor itself, a second, independent wall built out of different material.
A live-exploit test called Dirty Frag, run in 2026, made the gap concrete instead of theoretical. On the VM side, even configured with deliberately weak settings and an unpatched guest kernel, with the deck stacked in the attacker's favor, every attempt to break out and reach the host failed. That is the practical shape of the kernel-sharing boundary: a shared-kernel compromise ends at the host; a hypervisor-backed compromise ends at the guest, at least in that test.
gVisor sits in an interesting middle position here, and it's precise to say what it actually buys rather than treating it as a weaker cousin of a VM. The Sentry runs as an unprivileged, heavily seccomp-filtered userspace process, so a full compromise of the sandboxed application drops the attacker into that constrained process with a tiny allowed-syscall surface, not into kernel context directly. But the host kernel is still shared underneath it. A kernel vulnerability reachable through whatever syscalls the Sentry is allowed to make is, in principle, one exploit away from every tenant on that node. gVisor changes the odds of that exploit landing. It doesn't change the shape of what happens if it does.
Firecracker and Kata-backed VMs raise the bar to something categorically harder: a guest kernel escape followed by a hypervisor escape, against a VMM deliberately built small. KVM escapes are rare enough that Google funds a dedicated bounty program, kvmCTF, specifically to hunt for them, and hypervisor CVEs command some of the highest prices in the exploit market precisely because they're that hard to find. That scarcity is the practical measure of how tall the second wall actually is.
gVisor's tradeoffs: syscall overhead, density, and the host kernel's remaining reachability
gVisor's design rests on a specific bet: that intercepting every syscall in userspace can deliver meaningful isolation without paying the cold-start and memory cost that comes with booting a full virtual machine. It delivers meaningful isolation on its own engineering terms. It's a distinct engineering answer to the same problem, and it has a production track record to back it.
The mechanics: the Sentry reimplements the Linux ABI in userspace, and a second process, the Gofer, handles filesystem access from the host side, so even total compromise of the sandboxed application lands the attacker in a constrained userspace process rather than in the kernel itself. Operationally, this pays off in a way that matters a lot to platform teams: gVisor ships as an OCI runtime called runsc, sitting behind a Kubernetes RuntimeClass, so a pod opts in by setting one field and keeps every other normal pod behavior, the same images, the same volumes, the same kubectl workflow a team already runs.
Memory overhead runs about 18 to 50MB of Sentry process cost per sandbox, meaningful once you're running thousands of them but still well under what a QEMU-backed Kata deployment carries. The cost concentrates where you'd expect it to, given that every syscall is being intercepted: I/O-heavy operations and network paths take a bigger hit than CPU-bound computation, so an agent that's compiling code and streaming tool output back to a user sits closer to the expensive end of that range than one doing pure number-crunching.
There's a compatibility ceiling too, and it deserves to be named rather than glossed over. gVisor has implemented most of the Linux syscall surface, but every few months some team runs into a syscall it hasn't implemented, and a workload breaks. For an agent sandbox running arbitrary, user-supplied code, that is a real constraint on what you can promise will just work. It's a real constraint on what you can promise will just work.
None of this is theoretical at scale. Google runs gVisor as the isolation layer behind GKE Agent Sandbox, offered at no extra charge, specifically to let customers execute and test AI-generated code safely. Modal runs gVisor containers for isolation at the scale of a company that reached a unicorn valuation in a major Series C round in May 2026. That's a technology carrying real production load, proven at meaningful scale. It's carrying real production load for organizations with a lot to lose if the isolation story is wrong. gVisor works where nested virtualization is unavailable, making it a viable option on some cloud instance types that do not expose KVM to guests, though Kata Containers with PVM support has also recently gained the ability to run without hardware virtualization.
Firecracker's microVM model: the minimalism's benefits and omissions
Firecracker is a hypervisor in the narrowest, most literal sense: a userspace virtual machine monitor that uses KVM to boot a VM with its own guest kernel, built from the ground up for one workload shape, enormous numbers of short-lived VMs packed as densely as possible onto a single host.
The minimalism is the whole design philosophy, and it's a security decision as much as a performance one. Nothing else. The entire codebase is a fraction of what QEMU runs. That's a direct, mechanical link between doing less and being harder to break. Fly.io runs Machines on Firecracker and extended the model to Sprites (persistent microVMs aimed at agent workloads), entering that market in January 2026.
Snapshot-restore pushes the latency story further. Instead of cold-booting a fresh environment for every request, a fully initialized VM state gets baked into a snapshot once, and each subsequent restore maps that frozen state back in rather than repeating the whole boot sequence. Network throughput stays close to what the hardware can do natively too, with Firecracker's direct device model costing roughly an 8% penalty, a much smaller tax than the hit gVisor's syscall interception takes on network-heavy paths. Up to 150 microVM creations per second per host. The wmjournals.com 2026 engineering abstract reports that snapshot-restore cuts median cold start from 145ms to 28ms; PandaStack's production implementation clocks the full pipeline at a 179ms p50 and around 203ms p99.
None of that comes for free, and the gap that opens up is a real one. Firecracker's interface is a JSON API and a vsock channel. That's it. There's no pod semantic, no scheduler integration, no CNI plugin, nothing that looks like a Kubernetes runtime out of the box. Running it in production means building kernel image management, root filesystem handling, network configuration, the jailer security layer, and full VM lifecycle tooling, all from scratch. That's not a flaw in Firecracker's design; it's the design intent. Firecracker, gVisor, and Kata Containers are not the same kind of thing, and treating them as interchangeable options in a feature matrix produces a comparison that cannot settle anything. Most product teams are not in that business, and building it themselves is a serious undertaking rather than a footnote. Boot to userspace in under 125ms. Memory overhead per microVM under 5MiB.
The burden is tractable, though, and there's production evidence for it. One sandbox provider running every workload as a Firecracker microVM scaled from tens of thousands of sandboxes a month in early 2024 to millions a month by early 2025, a jump large enough to raise a significant Series A on the strength of it. That same provider later added pause-and-resume with filesystem and memory state kept intact, pausing in a few seconds per GiB of RAM and resuming faster still. Both are proof that raw Firecracker carries an operational cost that a team willing to own it can bear. Fewer lines means a smaller surface for hypervisor CVEs. Performance headline figures come from the Northflank January 2026 guide and OpenComputer guide.
Kata Containers: the orchestration layer above the hypervisor choice
Kata Containers is the answer to the gap the previous section just described: it's the orchestration layer above the hypervisor choice, not a fourth isolation technology competing with the other three. Pods still look like pods to the scheduler, the CNI, and kubectl, while each one actually runs inside a VM with its own guest kernel underneath, making microVM isolation behave like an ordinary container from Kubernetes' point of view.
The hypervisor choice inside Kata changes the numbers more than anything else about Kata. One platform builds its production microVM isolation primarily on Kata with Cloud Hypervisor, while also supporting Firecracker and gVisor as alternate backends, giving it the broadest spread of isolation options of anything named in this comparison.
Firecracker's sub-125ms boot time reflects raw hypervisor VM-boot performance, compared to roughly 300–500ms for QEMU-backed Kata. Firecracker backend: boot drops toward 125ms, overhead toward single-digit megabytes (roughly 5 MiB). Firecracker has no native virtio-fs or 9p support, though, so Kata-on-Firecracker requires the nydus snapshotter and a userspace nydusd daemon in the storage path, an extra dependency that Kata-on-QEMU does not need, the bex.co September 2026 analysis found.
That price buys something specific: kernel image provisioning, networking setup, and full VM lifecycle management, all handled, so a team gets hardware-level isolation without first turning itself into a virtualization company. Framed that way, "Firecracker versus Kata" isn't really a technology choice at all. It's a question of who owns the orchestration layer. Most organizations are better served running Firecracker through Kata Containers unless they already have a dedicated team built to own microVM orchestration directly, which makes Kata the sensible default for Kubernetes-native teams and raw Firecracker the right call mainly for teams building infrastructure products where owning that layer is the actual business. The hypervisor is pluggable: QEMU (best supported, best for GPUs and confidential computing), Cloud Hypervisor (Northflank's default, described as best performance), and Firecracker are all listed as supported backends in Kata's own documentation, the OpenComputer and Northflank guides note. QEMU backend: cold starts roughly 500ms–1s, guest overhead roughly 50–130MB (the heaviest footprint, buying the most mature device model), the bex.co September 2026 analysis found.


