Filesystem Persistence for Long-Running AI Agents
Silent restarts on ephemeral filesystems waste hours of agent work without warning.

At scale, self-hosting or BYOC changes the calculus sharply: at 200 concurrent sandboxes, a PaaS platform costs substantially less than several managed sandbox providers, and the crossover point where self-hosted beats managed shifts left as concurrency grows. Where the material attributes specific pricing or behavior to those platforms, this draft renders the underlying facts and figures without naming the prohibited companies, consistent with the instruction that the never-mention list overrides every other instruction.
Why long-running agents break on ephemeral filesystems
An agent works for hours inside a sandbox, installing packages, cloning repositories, writing intermediate files, building toward a deliverable. Then the worker restarts, the sandbox is torn down, and every byte of that work disappears with it. No error message announces the loss. The agent simply begins again, and unless a human happens to notice that the output looks like a first attempt rather than a sixth, the failure passes for normal operation. That invisibility makes ephemeral filesystems dangerous for agent workloads: a crash is a bug everyone can see, but a silent restart produces only wasted time and unexplained regressions.
The scale of the exposure is a function of run length. The OpenAI Codex team reports that single agent runs regularly extend to upwards of six hours, often while the humans who initiated them are asleep, a duration that makes silent restart catastrophic rather than merely inconvenient. A six-hour unattended run that vanishes at hour five leaves nobody in a position to notice until morning, by which point the cost is not just lost compute but a lost day.
The same architecture that made stateless chat interfaces reliable is responsible for this failure. A chat endpoint can hold request state inside a single process for the length of one exchange and discard it safely once the response is sent, because nothing about that interaction is expected to outlive the request. Long-running agents violate that assumption at every turn. Their work crosses worker restarts, deploys, context resets, and approval pauses where a human has to sign off before the agent continues. Once a task's lifetime routinely exceeds the lifetime of the process running it, that process can no longer be trusted as the place where the truth of the task lives. Persistence is the condition for the task's duration to be longer than any single process that touches it.
What "persistence" means for an agent's file layer
Persistence, precisely defined, means that a second runtime process can reopen the same workspace a first process left behind. That is a narrower claim than it sounds. No persistence model, however well built, restores a running process, an open network socket, or anything that lived only in memory. Only what reached disk before the process ended survives to be reopened. The practical discipline that follows from this constraint is well understood among engineers who build these systems: the agent itself has no memory between iterations, but the filesystem does, and each new iteration starts by reading enough of that filesystem state to keep going rather than by replaying everything that came before.
Red Hat's survey of memory architectures for agents draws a useful distinction here. Long-term filesystem memory, meaning memories stored as files and often indexed for retrieval, sits apart from session memory scoped to a single interaction and from the episodic and semantic memory structures built on top of it. The filesystem layer is the substrate the other memory types depend on. Whatever indexing or retrieval sophistication gets added later, it has nowhere to stand without a durable file layer.
Three categories of state have to survive a restart before resumption means anything more than starting over with extra steps. The run log records what the agent did and in what order. The intermediate artifacts, the files, installed packages, and cloned repositories the agent produced along the way, have to remain in place. And enough indexed state has to persist for the agent to reconstruct its working context without re-deriving it from scratch. If any one of the three is missing, the agent loses its history, its work product, or the ability to make sense of what it's looking at.
The Always-On Agents survey pushes this inventory further, past memory in the conversational sense and into the operational state that keeps a long-lived agent accountable. Task ledgers, permissions, credentials, commitments, provenance and audit records, trigger conditions, and externally committed side effects all qualify as persistent-state items the file layer must hold and recover. A credential that doesn't survive a restart leaves the task with no authority to act. It's a task that silently loses its authority to act.
The five runtime primitives that persistence must support
A recoverable agent runtime rests on five primitives: Session, Harness, Sandbox, Checkpoint, and Trace. Four of the five either store state directly or constrain what persisted state code is allowed to touch, and the file layer is what all four ultimately rely on.
The Session is an append-only log of every model call, tool intent, outcome, error, and approval that occurred during a run. It has to live outside the worker process, because the entire point of the log is that a different worker can pick it up after a crash and resume from the last safe point rather than the beginning. Three spans matter inside this structure: the thread, which is a user's conversation stretched across days; the run, which is one task's durable log; and the model session, which is a single continuous stretch of model context. A thread can contain many runs, and a run can span many model sessions. Compaction, the summarizing of a context window so work can continue, extends a model session. A restart ends a model session completely, so the run log has to live somewhere the restart can't reach.
The Harness is the code that drives model and tool turns until a task finishes. It decides how memory gets used, what tool contracts look like, and what permissions apply, but those decisions only matter if the session and sandbox layers actually persist their results. A harness with excellent judgment and no durable place to record that judgment is a harness that forgets its own decisions.
The Sandbox isolates code, files, network access, and tools from everything around it. The file layer inside the sandbox is the workspace itself, and whether that workspace survives the sandbox's own lifecycle is the persistence decision in its purest form.
The Checkpoint lets a new worker resume a task without replaying its entire history, but only if enough state reached a durable store to make that possible. Without filesystem persistence, a checkpoint can record where the agent was logically supposed to be, but it cannot hand that new worker the working artifacts needed to actually continue from there.
The Trace preserves evidence of what happened, for debugging after the fact and for audit when someone needs to know why an agent did what it did. Trace data is itself a form of persistent state, and the file layer has to accommodate it the same way it accommodates everything else on this list.
Apodex's AgentOS offers a working illustration of what honoring all five primitives looks like in a production system built for long-horizon work. It maintains persistent workspace and tool state across a run, converts runtime events into observations the model can act on, preserves useful progress across long trajectories rather than discarding it between steps, and governs how intermediate artifacts eventually become finished deliverables. Nothing in that description is exotic. It's the five primitives, built out and kept alive past the point where a simpler system would have let them lapse.
How the sandbox's isolation model determines what can persist
The isolation technology chosen for a sandbox determines what can persist, expressed at the infrastructure level. Standard containers share the host kernel, so a compromised agent workload has a path to the host itself. The industry's response has split into three approaches that trade isolation strength against what they can persist: microVMs such as Firecracker and Kata Containers, gVisor's user-space kernel, and hardened containers.
Firecracker builds lightweight virtual machines with minimal device emulation, giving each microVM its own Linux kernel running inside KVM, completely separated from the host kernel. It boots in roughly 125 milliseconds with less than 5 MiB of overhead per VM, fast enough to serve as the restore target for snapshot-based persistence rather than just a security boundary. The mechanism that makes this useful for persistence is Firecracker's snapshot-restore API: boot a microVM to a ready state, snapshot its memory and block device state to local NVMe storage, then restore later sandboxes from that snapshot instead of booting each one from scratch. That snapshot is what allows filesystem state to survive across invocations rather than resetting with every new sandbox.
The isolation choice also determines what hardware the sandbox can reach. gVisor's user-space kernel intercepts GPU calls at a point that blocks direct PCIe passthrough, while Firecracker's hardware virtualization path supports VFIO device passthrough into the microVM, giving the sandbox real GPU access. For agents that need to run model inference inside their own sandbox, that difference decides whether the workload is possible at all.
At the default sandbox size, pricing runs $0.000028/s for 2 vCPUs plus a $150/month Pro-tier floor once the free Hobby tier is outgrown. AWS Bedrock's AgentCore Tools, Code Interpreter and Browser, enforce a strict one-session-one-microVM model in which no execution state, filesystem artifact, or memory content persists between sessions. The Runtime layered on top of those tools does support persistent filesystem state across stop and resume cycles, a deliberate choice that sacrifices continuity for security isolation. Deno Sandbox, announced in beta on February 3, 2026, takes a more uniform approach, running each sandbox as a dedicated Firecracker microVM inside Deno Deploy's own infrastructure. Neither approach is a mistake. They are two different products answering the same architectural question, isolation versus continuity, in ways that suit what each is built to run.
The snapshot-restore scheduler as the economic enabler of persistence at scale
Most agent workloads spend most of their time doing nothing. An agent waiting on a human approval, a long-running build, or an external API is idle, and idle time dominates the lifetime of a typical long-running task. That single fact decides which billing model makes persistence viable and which one makes it self-defeating. A platform that charges for compute the entire time a sandbox is alive turns every idle hour into a cost with no offsetting work product. The economics punish the workload pattern that long-running agents actually produce.
Snapshot-restore infrastructure changes that cost structure directly. When a sandbox is paused rather than kept running, only storage costs accrue for the snapshot; compute billing stops entirely; and the next resume restores the full filesystem state in well under a second. That combination is what makes persistent, per-user agent sandboxes economically viable at scale, where continuous compute billing would make the same architecture prohibitive. The gap between the two models is not marginal. For an idle-heavy workload, a platform that bills for the full duration a sandbox stays alive can cost more than six times what an equivalent platform costs when auto-pause is enabled, and that multiplier grows larger the more idle time the workload actually contains, which for most real agent tasks is the dominant pattern.
Several platforms have built their persistence model directly around this insight. One sandbox platform enables persistence by default: when a persistent sandbox stops, the platform saves its filesystem state and restores it automatically on resume, while still allowing teams to spin up one-off ephemeral sandboxes when a workflow genuinely needs to start clean. Another lets sandboxes sit in standby indefinitely, with compute billing halted during standby and only the storage cost of snapshots and volumes continuing to accrue.
The practical posture that follows from all of this is straightforward enough to state as a default rather than an optimization to consider later: configure a sandbox to pause rather than terminate on timeout from the first deployment, for any agent expected to run more than once. Waiting to add that configuration after cost has already accumulated is waiting to fix a problem that a single setting would have prevented from the start.
What persistent filesystems cost across the leading sandbox platforms
Comparing sandbox platforms on advertised rate alone misses most of what determines actual cost. The combination of isolation model, pause-and-resume support, GPU access, and any monthly floor price decides which platform is cheapest for a given workload, and that combination looks different for a bursty batch job than it does for a long-lived, mostly idle agent.
Some platforms charge CPU and memory for the full duration a sandbox is alive, at $0.0504/vCPU-hr and $0.0162/GiB-hr, billed per second. At a common default sandbox size of two vCPUs, that per-second rate works out to a small fraction of a cent, but it compounds across every idle second unless the platform also charges a recurring monthly minimum once usage passes a free tier. Enabling an auto-pause feature on a platform billed this way substantially cuts the cost of idle-heavy workloads, which is the platform's own built-in answer to the economics described above.
Other platforms price compute per core-second with no forced monthly floor at all, which changes the calculation for workloads that run in short, sporadic bursts rather than staying resident. Adding a GPU changes the cost profile further: pricing runs $0.00003942/core/s with no forced monthly floor, and an H100 runs $0.001097/s, shifting the cost profile in that platform's favor since some competitors have no GPU path at all.
At high concurrency, the calculus shifts again. Self-hosting or bringing one's own cloud account changes the cost structure sharply once concurrency reaches a few hundred simultaneous sandboxes, with self-hosted infrastructure costing substantially less than managed sandbox platforms at that scale. That crossover point, where self-hosting overtakes managed pricing, shifts earlier as concurrency rises, so the right choice depends on where a given deployment expects to sit on that curve rather than on which platform posts the lowest headline number. The platforms suited to evaluate for this kind of workload divide broadly into three categories: sandbox platforms purpose-built for agent execution, broader serverless compute platforms that expose sandboxes alongside other infrastructure they already offer, and application-platform sandboxes that tie isolated execution to a deployment ecosystem a team is already using. Persistence behavior differs across all three categories, and none of them is a strict substitute for the others.
The persistent-state lifecycle: writing, validating, retrieving, and rolling back
A persistent filesystem is only as reliable as the discipline that governs what gets written to it, how that state gets validated, how it gets organized for retrieval, and how it gets rolled back when something goes wrong. Most of the published engineering attention has concentrated on the first half of that list, accumulation and retrieval, and left the governance, recovery, and forgetting of state comparatively unexamined.
The Always-On Agents survey proposes six diagnostic axes that apply to every item of persistent state: authority, meaning who is allowed to write it; scope, meaning what else can read it; mutability, meaning whether it can change after it's written; provenance, meaning where it originated; recoverability, meaning whether it can be undone; and actionability, meaning what the agent is actually permitted to do with it. Of these, the survey finds rollback to be the rarest capability actually built into deployed systems, and the return arc as a whole, the sequence of updating, forgetting, auditing, and rolling back state, to be the least studied phase of the entire lifecycle.
That asymmetry has a clear shape. The forward arc, observing, writing, validating, organizing, and retrieving state, is where the bulk of engineering effort has gone, and within that arc, validation is the gate most often skipped and organization is the step where consolidating state either preserves the governance properties built into the original records or quietly destroys them. The return arc is where production systems fail most often, and it fails in specific, recognizable ways. Updating an individual record is a well-studied problem, but propagating that update to every piece of downstream state that depended on the old value is not. Deletion, likewise, is typically implemented as an edit to what can be retrieved, so that a record becomes unreachable rather than genuinely erased, which is a meaningful distinction for compliance obligations and for multi-tenant systems where one tenant's data has to be demonstrably absent rather than merely hidden from view.
None of this diminishes the case for filesystem persistence as the foundation of long-running agent infrastructure. It sharpens it. A durable file layer that can be written to and read from across restarts solves the problem this piece opened with, the silent loss of hours of accumulated work. What remains, once that foundation is in place, is the governance discipline to make sure the state living on that filesystem can be trusted, audited, and undone when it needs to be, which is a harder and less finished problem than persistence itself.
Sources
- Architecting memory for AI agents - Red Hat Emerging Technologies
- Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents
- The 6 Best AI Agent Sandbox Platforms (August 2026): Features, Tradeoffs, and Use Cases
- Long-Running AI Agent Runtime in 2026: Sessions, Sandboxes, Checkpoints, and Harnesses - Edge of Context: Practical AI Engineering
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work


