Est.

Agent State Recovery After Unexpected VM Termination

How to design agents that survive infrastructure failures and VM crashes.

Staff Writer · · 11 min read
Cover illustration for “Agent State Recovery After Unexpected VM Termination”
Persistent Agents · October 7, 2026 · 11 min read · 2,432 words

An agent running a multi-hour task inside a worker process will, at some point, lose that process before the task finishes. This is a structural mismatch between how long these tasks run and how long any single VM or process can be expected to stay alive. Hosts reclaim memory under pressure, killing processes when it is oversubscribed and finally claimed in full. Processes crash mid-tool-call. Infrastructure restarts land wherever they land, often in the middle of an agent writing to a database, sending an email, or calling a payment API. None of this is hypothetical, and none of it waits for a convenient moment in the task.

The gap between what organizations expect from their agent deployments and what they can actually handle is well documented. Cohesity's research found that a majority of organizations are not well prepared to detect, contain, and recover from unintended or incorrect actions taken by AI agents, and that finding doesn't even account for the added complication of a VM disappearing mid-run. Recovering from an agent's bad decision at least starts from a complete record of what it did, so adding VM loss to that picture makes the preparedness gap worse. Recovering from a terminated VM means reconstructing what happened from whatever was saved before the termination, which, in most deployments, is close to nothing.

The conclusion that follows is specific. An agent that survives VM termination is designed, from its first architectural decision, around a different set of primitives than the ones most frameworks ship with by default. Survival is a property of the design, not a patch applied after the first outage.

Why the context window is not the agent's state

The architectural error behind most of this fragility is simple to state and easy to miss in practice: treating the provider's conversation object, response ID, or in-memory message list as if it were the full record of what the agent has done. It feels like state, because it contains the history of the exchange. It is not state in any sense that survives a process restart, because the moment the worker process dies, that history goes with it.

Large language models hold nothing between API calls. Every call is stateless from the model's side, so memory, tool outputs, credentials, approval records, and task status all have to live somewhere else, in an explicit layer outside the model, or they vanish the instant the process that was tracking them disappears. It's the default behavior of most agent frameworks unless someone explicitly builds the external layer to override it.

The real state of a production agent is much larger than its conversation history. It includes task status, the inputs and outputs of every tool call, records of user approvals, files the agent has touched, evidence it has retrieved, intermediate summaries it has written, retry counters, budgets, leases, timestamps, and records of which external side effects have already happened. Any one of these living only inside a worker's memory is a data-loss bug waiting for a restart to expose it. Multiple agent runs touching the same workspace files add a second failure mode on top of the first: one run's mutations can quietly corrupt the assumptions another run is relying on, so even a process that survives can wake up holding bad state.

The design principle that follows is narrow but consequential. Preserve the state of the work, not the text of the conversation, keeping that state outside the model provider. Done this way, a provider outage or a VM termination changes only which inference call has to be retried. The task record itself stays intact.

The five architectural primitives that make an agent recoverable

Diagram: The Five Primitives of a Recoverable Agent. Visualizes: Show the five architectural primitives that make an agent recoverable, in the order they chain together during an actual recovery sequence: (1) External, append-only state, (2)…

Once state lives outside the model, the rest of the system has to be organized around it. The field has converged on a runtime model in which the infrastructure layer, not the model, owns session state, tool execution, checkpoints, policy checks, secrets, traces, cost limits, and deployment shape. The model's job shrinks to choosing the next action, given everything the infrastructure hands it.

Five primitives make that model work, functioning together as a system. External, append-only state keeps every meaningful completed step recorded outside the model provider before the run advances, so there is always an authoritative record to recover from. Durable state boundaries mean tool results get written to storage before the run moves to the next step, so a crash between steps loses nothing that already finished. Checkpoint validation adds a check on top of that record: a saved state is not automatically safe to resume from, so each checkpoint needs evidence attached that confirms it's a legitimate place to continue. Idempotent tool calls ensure that replaying a step, which recovery sometimes requires, doesn't double-charge a card, send a duplicate email, or close a ticket twice. Fast-restore execution environments make all of the above affordable, by reconstituting an agent's working context, its mounted workspace, its progress files, its last checkpoint, in well under a second.

These primitives chain together in the actual recovery sequence. A new worker mounts the workspace from its last known state, reads whatever progress files the previous attempt left behind, and loads the last checkpoint from the database. That sequence is the harness handing a replacement worker everything the dead one was holding in memory, reconstructed from disk.

What checkpoint validation requires

Checkpoint validation is the primitive most teams get wrong, because the naive version of it fails in both directions. Restarting a task from scratch after a crash throws away work that already passed validation. Restoring the most recent checkpoint without checking it can carry an earlier, undetected error forward into everything that happens next.

A 2026 arXiv paper on recoverability in long-horizon AI agents formalizes this problem directly. Its central claim is that a state being saved does not mean it remains suitable for continuation. The system has to explicitly select a supported starting point and a permitted recovery action, or withhold automatic continuation when it cannot.

The paper gives this shape through three concepts. A continuation source is the specific checkpoint chosen as the recovery point, and it is not necessarily the most recent one saved. A recovery route is the corrective action permitted from that source: retrying a read is usually fine, but retrying a tool call that may have already changed something in the outside world might not be. Continuation authority is permission granted at the moment of failure, based on evidence and policy, rather than a permanent label stamped onto a checkpoint once and trusted forever; that authority can expire.

The paper's own example makes the stakes concrete. Checkpoint A holds a document version that has already been approved. Checkpoint B holds a later, exploratory edit that failed a required check. If the system restores B and successfully repairs the failed edit, the final document looks fine. But a document that reads correctly at the end says nothing about whether the recovery started from a permitted point. Task success at the end hides a bad decision made in the middle.

That gap is the practical argument for treating checkpoint evaluation as its own step, separate from how faithfully a restore reproduces prior state and separate from whether the task eventually completes. The policy evidence that grants continuation authority needs to be held independently of the checkpoint itself, so it can expose a violation in how recovery happened even after the effect of that recovery is already sitting in the world.

Why idempotency is the hardest primitive to retrofit

Idempotency is hard to add after the fact because the failure it has to protect against is asymmetric. A crash before a tool call runs just loses the work the call would have done. A crash after the tool call runs, when the system then retries it on resume, duplicates a real action: the same email sent twice, the same card charged twice, the same ticket closed twice when it was already closed the first time.

Picture a payment tool call that posts a charge to a processor and then crashes before the result of that charge gets written to storage. On resume, a naive retry sees no record of success and calls the payment API again. The customer is charged twice, and the system has no way to know it happened, because nothing recorded that the first call ever completed.

A resilient architecture treats this as something to design for from the start, not something to patch in later. A stable ID gets assigned to every run. Every step carries a status. Every external side effect has an idempotency strategy attached to it before the system ever calls it in production. Every provider request can be retried or redirected without the system pretending the agent is starting the task over.

The distinction that matters here is between what's safe to repeat and what isn't. A model call is almost always safe to run again; nothing in the world changes because the model generated a response twice. A payment, an email, a ticket closure, a database mutation, or the consumption of an approval is a different category of action, and a generic retry wrapper at the routing layer has no way to tell the two apart. Persisting the tool's result before the run advances is what actually solves this: the checkpoint records not just that a step was attempted but that its result was received and stored, so a resumed run reads that stored result instead of calling the tool again.

This is why idempotency resists retrofitting. An agent built on a stateless framework, where tool calls fire and their results live only in the context window, has no ledger anywhere recording what already happened. Wrapping that agent in retry logic after the fact doesn't create the missing record; it just adds a layer that retries blindly, because there's nothing underneath it to consult.

How fast-restore execution environments make recovery economically viable

Checkpointing and idempotency solve the correctness side of recovery, but correctness alone doesn't make recovery something teams actually use. If resuming an agent from a checkpoint takes several seconds of cold-start time, or forces a team to pay for idle compute while that wait happens, operators default to full restarts instead and recovery stays theoretical.

Firecracker, the microVM technology, addresses this directly. Each microVM boots in well under a second, with minimal memory overhead, because Firecracker skips firmware entirely and boots a minimalist Linux kernel directly, with no BIOS and no POST sequence to wait through. Snapshot-restore lets a VM be frozen and brought back per request in tens of milliseconds, making it economically workable to spin up a fresh VM for every single invocation. Copy-on-write fork lets a running VM's memory get branched cheaply, which maps onto a pattern agent recovery actually needs: resuming five parallel fix attempts from the same validated checkpoint without paying to rebuild that checkpoint's environment five separate times.

A 2026 system called DeltaBox builds on this with a warm-template approach. Checkpointed processes get preserved as frozen templates, OS-level fork() enables low-millisecond restores from those templates, and a Network Proxy Daemon decouples the agent's LLM I/O from its address space, so the templates stay safely forkable without the agent's network state getting tangled up in the fork. Amazon Bedrock AgentCore Runtime applies the same underlying technology at the product layer, running each agent session in its own Firecracker microVM under a one-session-one-microVM architecture.

The isolation this provides and the recovery it enables are the same property seen from two angles. Containers and permission prompts run in userspace, the same layer the agent itself reasons and acts in, so a compromised or misbehaving agent can potentially subvert them. A microVM's isolation is enforced by hardware, underneath the layer the agent can reach. A recovered agent therefore resumes into an environment it couldn't have corrupted even if the run that crashed was already misbehaving.

The economics of idle agents and the snapshot-standby resolution

Agent workloads have a cost profile that punishes naive infrastructure. A large share of any given session is spent idle: waiting on an LLM response, waiting on a human approval, waiting on a downstream API to answer. Billing by the wall clock turns all of that waiting into pure cost, charged for compute that is doing nothing.

Snapshot-standby architecture is the structural fix. When an agent goes idle, its full execution state gets snapshotted to disk. When a new request arrives, that snapshot restores in tens of milliseconds. While paused, the only ongoing cost is the price of storing the snapshot, not the price of holding a VM awake and idle.

It's the same snapshot-restore primitive doing double duty, powering the billing trick and the recovery story alike. An agent that's economically parked on disk between requests is, mechanically, in exactly the state a crashed agent needs to be restored from: a validated checkpoint sitting in storage, ready to resume. What differs is only the trigger, idle timeout in one case, unexpected termination in the other.

Applied across a multi-tenant fleet, aggressive idle timeouts paired with a memory-suspend mode let each tenant's agent drop off the billing meter almost immediately when it's not doing anything, and come back in sub-second time when it's needed again. That is what makes deploying a dedicated agent per user, rather than a shared pool, viable at the scale a real product needs.

Durable-Execution Engines in Production Agent Stacks

Most production teams don't implement these five primitives from raw infrastructure. They wrap an existing agent framework in a durable-execution engine that owns checkpoint progression, retry of individual activities, and state persistence, so the person building the agent gets recoverability without writing the underlying machinery themselves.

Temporal's integration with the OpenAI Agents SDK reached General Availability on March 23, 2026. Its activity_as_tool() pattern converts any Temporal Activity directly into a tool the agent can call. If an agent calls a database_query tool that's backed, under the hood, by a Temporal Activity, a crash during that call doesn't lose the work: Temporal retries the activity on its own, and the agent picks back up from its last durable boundary.

The Vercel AI SDK integration applies the same underlying idea in TypeScript. LLM API calls are non-deterministic. They can't sit directly inside workflow logic the way a deterministic function can. The integration wraps generateText(), streamText(), and streamObject() in Activities behind the scenes, so the person writing the agent never has to manage that wrapping explicitly. The durable-execution layer handles it: the agent code reads as if it were calling the LLM directly, while the actual execution is checkpointed, retryable, and recoverable underneath.

Sources

  1. Recoverability as a System Primitive for Long-Horizon AI Agents
  2. Firecracker
  3. Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents
  4. AgentR A Stateful and Recovery-Aware Software Architecture for LLM-based Auditable Workflows
  5. Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model
  6. Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
  7. The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems
  8. Recoverable Execution for Long-Horizon LLM Agents

More in Persistent Agents