Cold Start Latency Benchmarks for Agent VM Runtimes
Micro-VMs with snapshot-restore beat containers for agent tool calls by orders of magnitude.

An agent that ran perfectly in staging starts to feel sluggish the moment it hits production traffic, and the model is rarely the reason. The reason is almost always cold starts, but not in the way most engineers have been trained to think about them. In an agent loop, tool calls run one after another inside a single interaction: the user sits and waits for the sum of every cold-start penalty across every tool call before any result comes back. That's a structurally different failure mode than the one most infrastructure teams have spent a decade optimizing around.
Serverless engineers know the old math. A cold start hits one request in isolation, the next request probably lands on a warm instance, and the penalty is a rounding error spread across a fleet. Agents break that assumption completely. The overhead is closer to total overhead = tool calls times per-call cold-start probability times average cold-start duration, and even a modest per-call rate turns into something a user can feel once you multiply it across enough steps. Academic benchmarks cited in industry analysis put average tool calls per task at roughly 30 for Claude Opus and roughly 44 for GPT-5.4 running at high reasoning settings. A cold start of even one second per call can eat the entire interaction budget before the agent has done anything useful.
This is why so many teams misdiagnose the problem. Most observability stacks measure end-to-end response time and pin the blame on model inference, since that's the biggest number on the dashboard and the most familiar villain. Splitting the trace into per-tool-call spans, using something like OpenTelemetry to separate "environment ready" from "execution complete," reveals a different picture: the environment-ready span often outweighs the execution span by several multiples on short-running tools. The model was never slow. The sandbox was.
The substrate taxonomy: what determines a cold-start number
Every cold-start number traces back to three variables: how the platform provisions the runtime, what device model it exposes to the guest, and where the isolation boundary sits. Those three factors explain why cold-start figures across substrates span more than two orders of magnitude, from tens of milliseconds to tens of seconds, for what looks on paper like the same job.
Four substrate classes cover essentially everything an agent workload might run on. Cloud VMs are one extreme: strong isolation, the highest provisioning cost, and cold starts measured in tens of seconds, because they were never designed to be created per request or per agent step in the first place. Kubernetes-orchestrated containers add scheduling overhead on top of container startup itself, and that overhead saturates fast: Dirigent research shows Knative and OpenWhisk hitting a ceiling below roughly two cold starts per second, driven by CPU pressure on the Kubernetes API server. Hosted sandboxes abstract the substrate away entirely, and because of that, the resulting cold-start numbers vary widely depending on what's actually running underneath the managed layer. Micro-VMs, Firecracker and its derivatives chief among them, strip out the device model almost entirely: no USB, no PCI bus, no BIOS, none of the legacy hardware emulation a normal VM carries. That's an architectural choice, not a shortcut, and it's the single biggest reason micro-VMs win on both attack surface and per-VM memory footprint at scale.
Firecracker is the open-source VMM behind AWS Lambda, Fly.io, and E2B, and its reduced device model is the specific design decision that makes high-density, low-latency VM creation possible in production. The asymmetry between substrates is load-sensitive, too. For a trivial task, the gap between the fastest hosted sandbox and a cloud VM can run past two orders of magnitude, but as the workload gets heavier, execution time increasingly absorbs the boot cost and the gap narrows. Agents tend to run many short tool calls rather than a few long ones, so they feel substrate choice more acutely than almost any other workload class.
Snapshot-restore is the mechanism that turns Firecracker from "fast micro-VM" into "fast enough for agent loops." Instead of booting a kernel fresh on every create, the platform bakes a snapshot of a VM that's already booted: kernel running, guest agent running, network stack initialized, all of it frozen at a known-good point. Restoring from that snapshot on demand replaces a boot sequence with something closer to a memory copy.
PandaStack's production deployment shows what that looks like in practice. Every sandbox create restores a Firecracker snapshot rather than booting cold; the restore step itself takes roughly 49 milliseconds, and the p50 and p99 for the full end-to-end create are both in the low hundreds of milliseconds. The only place the full cold-boot cost gets paid is the first time a template boots, before any snapshot exists to restore from, and that cost is absorbed once and amortized across every subsequent create.
Latency numbers only tell half the story, though. Dirigent, running on Firecracker micro-VMs, reaches a peak throughput of 2500 cold starts per second while control-plane CPU utilization stays comfortably below saturation, which is the evidence that this architecture scales horizontally instead of choking on its own control plane. Set that next to Knative and OpenWhisk saturating below two cold starts per second: the contrast is a difference in what the substrate was built to do, not a tuning difference between two similarly designed systems. It's a difference in what the substrate was built to do. Firecracker's numbers, in other words, are a production floor that other substrates have to be measured against. They're a production floor that other substrates have to be measured against.
How snapshot-restore and delta-checkpointing extend what micro-VMs can do for agents specifically
Agents don't execute in a straight line the way a typical serverless function does. They branch, they backtrack, they fan out across multiple candidate paths and abandon most of them, and every one of those moves requires the platform to checkpoint and roll back complete sandbox state fast enough that the branching doesn't itself become the bottleneck.
Conventional checkpoint-restore duplicates the entire VM state on every checkpoint. For a sandbox of any real size, that means hundreds of milliseconds to multiple seconds of overhead per checkpoint, and under a fixed time budget, that overhead becomes a hard ceiling on how many branches an agent can actually explore. Research on a system called DeltaBox starts from a specific observation: consecutive checkpoints taken by an agent tend to look almost identical to each other, with only small incremental changes between one step and the next. Rather than duplicating the full VM state every time, DeltaBox duplicates only the delta. Evaluations on SWE-bench and reinforcement-learning micro-benchmarks show agents exploring substantially more of the search tree under the same fixed time budget as a result. DeltaBox is a research result, not a shipping product; it points at where the field is heading rather than something a developer can deploy today.
Something close to the production version of this idea already exists at the hosted-sandbox layer, though. E2B's pause and resume API lets a running sandbox, filesystem, memory state, running processes, loaded variables and all, get paused and picked back up later. Pausing takes time roughly proportional to how much RAM the sandbox is using, resuming takes a matter of seconds, and a paused sandbox just sits there indefinitely until someone explicitly kills it. That's a meaningfully different guarantee than restarting from a fresh boot, because it preserves exactly the kind of in-memory state an agent's next step might depend on.
Snapshot-restore isn't only a latency optimization bolted onto an existing architecture. It decides what operational posture a platform can actually offer, whether that's always-on, scale-to-zero, or hibernate-and-wake, and it decides whether branching agent search is something the platform can support at all, rather than something bolted on top of an architecture that wasn't built for it.
What the hosted-sandbox benchmarks show across E2B, Modal, Daytona, Vercel Sandbox, and Fly Machines
Cold-start numbers across hosted sandbox platforms vary across a range spanning more than two orders of magnitude, making a choice made without measuring functionally a coin flip. That's not a knock on any individual platform. It's a description of how differently "sandbox" gets implemented under the hood depending on which vendor is running it.
The test ran on a 12th Gen Intel i7-12700KF with dual-channel DDR4, Node.js v24.13.0, and Linux 6.1.0 on Debian, x86_64, with enough iterations that the p95 and p99 figures carry real statistical weight. The workload itself was deliberately minimal, a sleep task on an idle Node.js process, which measures the platform's floor rather than any application's actual initialization time.
AWS Lambda's 2026 cold-start figures, drawn from production workloads rather than synthetic benchmarks, give a useful second reference point for anyone comparing runtimes rather than sandbox platforms specifically. Rust on arm64 posts a p50 of 14.1 milliseconds and a p99 of 31.9 milliseconds, coming in faster than the blog-claimed range that preceded the measurement. Rust on x86_64 runs at a p50 of 17.0 milliseconds and a p99 of 29.3 milliseconds. Go on arm64 runs at a p50 of 45.0 milliseconds and a p99 of 61.2 milliseconds, again beating its earlier claimed range, which was somewhat higher. Python 3.13 runs at a p50 in the low hundreds of milliseconds with a p99 approaching a full second, and Node.js 22 tracks closely behind it, with a similar p50 and a p99 in the same neighborhood. Running any of these on arm64 Graviton instances shaves a meaningful chunk off the number, whichever runtime is in play.
The mechanism that produces those Lambda numbers matters as much as the numbers themselves. Lambda moved onto Firecracker micro-VMs in 2018, and VPC cold-start overhead, once a multi-second tax on every networked function, now approaches zero for most workloads. SnapStart extends the same snapshot-restore logic covered earlier: AWS snapshots the initialized Firecracker micro-VM right after its INIT code finishes running, then caches that snapshot across three tiers, one on the worker itself, one at the placement-group level, and one in S3 at the regional level. Python and.NET support for SnapStart went generally available in late 2024. Java 21 without SnapStart shows a p50 of 2–5 s and a p99 of 6–10 s. Java 21 with SnapStart shows a p50 in the low hundreds of milliseconds and a p99 in the low hundreds of milliseconds.
Agent-specific runtimes are starting to reflect this thinking directly rather than inheriting it from general-purpose serverless. AWS's AgentCore Runtime v2, which shipped on September 18, 2026, uses micro-VM-based isolation paired with elastic memory management, and its p75 cold-start figures now run a few seconds, down from as long as 30 seconds on the original AgentCore runtime. AWS built that improvement around an explicit premise: long-running, stateful agents are a distinct workload class, not a variant of a web request, and the v2 design reflects that directly. Firecracker's snapshot-restore approach beats that categorically for any workload creating sandboxes repeatedly. agentOS (Rivet) benchmarks are measured against E2B as the fastest mainstream sandbox, per Rivet's documentation. agentOS reports a cold start p50 of 4.8 ms, p95 of 5.6 ms, and p99 of 6.1 ms.
Confidential VMs and the additional latency cost of hardware-backed memory encryption
Confidential VMs add a layer of hardware-backed memory encryption, through Intel TDX, AMD SEV-SNP, or ARM CCA, that carries its own cold-start and warm-start cost. That cost sits on top of the baseline substrate costs already covered here, not instead of them.
Warm-start overhead in CVMs is workload-dependent: workloads with frequent VM exits, particularly idle transitions, suffer substantial slowdowns, and these slowdowns are further amplified by the common serverless practice of coupling vCPU allocation to memory size, which exposes more vCPUs than some functions use effectively.
None of this argues against CVMs where the threat model calls for them, healthcare, finance, and privacy-preserving inference being the obvious cases. It argues for benchmarking the specific workload inside a CVM context rather than assuming general-purpose overhead figures from the literature will map cleanly onto any particular agent's access pattern. Memory encryption disables host-level cross-VM page deduplication, which reduces warm-container capacity under a fixed memory budget and limits the density optimizations that make micro-VM economics work at scale, per the NJIT/Hofstra study (arxiv 2609.04478). For cold starts specifically, the study focuses on container creation within CVMs (a common and major contributor to startup latency in this deployment model) and shows that selectively relaxing certain isolation mechanisms when multiple containers of the same function are consolidated within one CVM can substantially reduce startup overhead, per the same paper.
Cold-start latency is an isolation decision, not a performance tuning knob
Everything above points to the same underlying fact: whether disposable, per-request isolation is something a platform can actually afford to offer depends on the cold-start number that platform hits, and if the number doesn't support it, the platform quietly falls back to a weaker lifetime model instead.
At a p50 of 179 milliseconds, a platform can afford to throw away the environment after every request, every agent step, every job, and start clean each time. No state survives long enough to leak from one use to the next, so cross-request contamination becomes structurally impossible rather than merely a risk to manage.
That's the real reason cold-start benchmarks deserve close reading rather than a glance at a headline number. A fast cold start isn't a nice-to-have optimization sitting on top of an architecture. It's the precondition for choosing the strongest isolation boundary available, per request, per step, per branch of a search tree, without paying for that safety in latency the user can feel. Slow cold starts don't just cost time. They quietly narrow which architectures are even on the table, and the only way to know which side of that line a given substrate falls on is to measure it directly, against the specific shape of the agent workload actually running on it.


