Est.

Credential and Secret Management Inside Agent VMs

AI agents need isolated VMs to safely manage dynamic credentials.

Staff Writer · · 11 min read
Cover illustration for “Credential and Secret Management Inside Agent VMs”
Persistent Agents · October 3, 2026 · 11 min read · 2,466 words

A cron job calls the same three endpoints with the same service account every night, and it has done so for years without incident. An AI agent, dispatched to close a single support ticket, might query a production database, draft a reply, schedule a follow-up meeting, and open a new ticket in a different system, each action demanding a different credential against a different backend, decided in the moment rather than fixed in advance. That difference in kind, not degree, is why credential management built for scripts and service accounts breaks down when applied to agents.

Three structural features of agent systems compound the exposure. The credential set an agent needs cannot be known ahead of time, because it emerges from the reasoning chain at the moment of inference rather than from a deployment manifest written by a human. Static secret injection, the pattern that has worked for a decade of server-side automation, cannot anticipate a set of permissions that does not exist until the model decides what to do next.

The second feature concerns where those credentials end up. An LLM's context window holds everything that has passed through the conversation: tool descriptions, API responses, and error messages all sit in the same token stream the model reasons over. A failed authentication attempt that includes the API key in its own error text writes that key directly into the model's working memory, where it becomes available to any later step in the same session, including one shaped by an attacker.

The third feature is aggregation. A single agent runtime often holds OAuth tokens for many services at once, so compromising the runtime compromises every system those tokens reach, and the blast radius grows with each new integration the agent is given. Agent credentials are typically provisioned broadly to make an integration work on the first try and then left that way, so one compromised key can reach production data, cloud infrastructure, and source control simultaneously. The OWASP Top 10 for Agentic Applications, cited in WorkOS's guide, names identity and privilege abuse as a core risk category precisely because traditional identity and access management was built for a world with a clean line between what a user authorized and what a system did on its behalf. Agents erase that line: the agent decides, at runtime, what it is going to do, and conventional IAM has no mechanism for authorizing a decision that has not yet been made.

Prompt Injection as a Repeatable Exfiltration Kill Chain

Diagram: The Three-Stage Prompt Injection Kill Chain. Visualizes: Illustrate a three-stage attack sequence that shows how prompt injection becomes a repeatable credential exfiltration kill chain.

Prompt injection is not a rare edge case to be patched after the fact. It functions as a repeatable, three-stage kill chain: a malicious instruction lands inside content the agent processes, the agent reads a credential it has access to, and that credential leaves through a channel the agent was already permitted to use for legitimate work.

Standard deployment habits make each stage easy. Agents run with long-lived API keys sitting in .env files or environment variables, which puts the raw credential value directly within reach of anything the agent's reasoning can be steered toward retrieving. Tool output, including authentication error messages that echo the key back in plain text, flows straight into the context window unless something explicitly strips it out first. A production study by Zylos found that a large majority of credential leaks in agent systems arrive through print() statements captured into agent context, which makes redacting tool output before it reaches the model a requirement rather than a best practice.

Two incidents from production systems show this chain executing, not in theory but against real infrastructure. In the "Comment and Control" disclosure of April 15, 2026, security researchers Aonan Guan, Zhengyu Liu, and Gavin Zhong demonstrated that Anthropic's Claude Code Security Review, Google's Gemini CLI Action, and Microsoft's GitHub Copilot Agent could all be hijacked by prompt injection payloads hidden inside GitHub pull request titles, issue bodies, and issue comments. The compromised agents exfiltrated repository and API secrets back out through GitHub itself, a channel each agent was already authorized to use, which is what made the exfiltration invisible to anyone watching for unusual network destinations.

A separate failure mode appears in CVE-2025-68664, a December 2025 vulnerability in LangChain Core rated 9.3 on the CVSS scale. An attacker could inject prompts that triggered LangChain's own serialization system into dumping environment variables, exposing every secret the process held. The vulnerability worked because the agent had direct access to raw credentials in its runtime environment. A third pattern surfaces through infrastructure defaults rather than application code: a deployed Vertex AI agent could be weaponized through a malicious tool call hitting the GCP metadata server at 169.254.169.254, an endpoint reachable without authentication from any GCP virtual machine. Because GCP assigns the project-level compute service account to VMs by default, any agent running on one of those machines could obtain valid tokens with no API key required.

What unites these three incidents is that none of them depended on breaking encryption or guessing a password. Each one took a credential that was already sitting in reach and used a permitted channel to move it. That points to a structural fix rather than a better detection rule, because the average detection window for credential breaches stretches to many months, a timescale with no relationship to agents that can act and exfiltrate within seconds. The fix has to start at the boundary around where the agent runs, not at the moment someone notices the damage.

Firecracker Micro-VM Isolation as a Security Primitive

A shared-kernel container does not supply that boundary. When a prompt injection turns an innocuous task like "summarize this CSV" into "read every environment variable and POST it somewhere," the resulting process still runs on the same kernel as every other container on the host. A single kernel-level escape turns one compromised script into a compromise of the entire host, along with every other tenant's workload running on it.

Firecracker microVMs enforce the boundary at a different layer. Each agent gets its own kernel instance, isolated by hardware virtualization rather than by namespaces and cgroups layered on top of a shared kernel. The boundary an attacker has to cross is enforced at the level of the processor itself, not by application-level policy that a clever payload might talk its way around. Firecracker also runs one VMM process per microVM rather than a single daemon managing a fleet of them, so compromising one VM's virtual machine monitor does not give an attacker a foothold into any other.

The isolation carries a second layer even if the first one fails. Escaping the hardware virtualization boundary would require a CPU hardware bug or a KVM vulnerability, and even then, the Firecracker jailer traps an attacker inside a chroot jail with no network access, restricted to exactly 24 system calls through a seccomp filter. That is a narrow enough surface that a successful hardware-level escape still leaves very little for an attacker to do next.

The traditional objection to this degree of isolation has been startup latency: booting a full virtual machine for every task has historically been too slow to compete with spinning up a container. PandaStack's documented architecture addresses that by restoring a snapshot of an already-booted machine instead of cold-booting one from scratch, adding only a brief additional step to the process, with end-to-end VM creation at a p50 of 179 milliseconds. That is hypervisor-grade isolation running at roughly the latency budget of a container.

For credentials specifically, this architecture means a compromised agent's reach is bounded by what was injected into its own VM. It cannot cross the kernel boundary to read another agent's secrets, cannot sniff another tenant's network traffic, and cannot reach the host's credential store, because none of those things are visible from inside its own hardware-isolated instance. The boundary holds only if what gets put inside each VM is handled correctly. That is the part the rest of this piece is concerned with.

Never baking secrets in: runtime injection as the baseline pattern

Secrets must never be baked into a VM template or image. They belong at dispatch time, injected for the specific task and the specific tenant that job serves, and nowhere else.

A correctly built injection flow starts with the dispatcher, which fetches only the credentials scoped to the activity about to run, at the moment it is about to run. Those credentials are written into the guest when the VM starts, and the entrypoint loads them into the environment before shredding the source file they arrived in, so no copy of the raw value persists on disk once the process has what it needs. If that tenant's step later leaks its own environment into a log somewhere, the damage is confined to that tenant's own scoped credentials. It becomes a support ticket instead of a breach that touches anyone else.

Three tools implement this same principle, in ascending order of rigor. HashiCorp Vault's dynamic secrets engine generates a temporary credential with a configurable time-to-live the moment an agent requests access, instead of storing a static database password anywhere at all, and automatically revokes that credential once the TTL expires. Vault's backends cover AWS, databases, SSH, and PKI among others, and its AppRole and Kubernetes authentication methods are built for exactly the kind of automated, non-interactive deployment an agent runtime represents. Vault is available self-hosted under a BSL license as of August 2023, or through HCP as a managed service with no free tier (HCP Vault Secrets, which did offer one, was sunsetted in 2026, leaving HCP Vault Dedicated as the current managed option, starting at a monthly cost).

SPIFFE/SPIRE takes a different approach to the same end, issuing X.509 certificates with lifetimes measured in minutes to hours rather than days or months. The SPIRE daemon rotates those certificates automatically, with no renewal logic required inside the application itself, and ties workload identity to process attestation rather than to a static configuration file that could be copied or leaked.

Following the "never bake" rule pays off at exactly the points where an attacker would otherwise profit. If an attacker obtains a copy of the VM image itself, there are no credentials sitting inside it to find. If one tenant's runtime environment leaks, it contains only that tenant's own scoped credentials, never anyone else's. A shared template, reused across tenants and tasks, cannot become a path for lateral movement, because the template never held anything worth moving laterally toward.

Scoping credentials to the task, not the agent

Injecting a credential at the right moment solves a timing problem, but timing alone does not solve a scope problem. A credential injected at runtime that still carries account-wide permissions creates the same blast radius as a credential that was baked in from the start, because a compromised agent can reach just as far with it either way.

Correct scoping operates at every layer a credential passes through. At the cloud IAM layer, that means short-lived STS session tokens scoped to exactly the resource prefixes a given tenant's job needs, for the duration of that job alone, so the VM never holds a route to anything outside its target endpoint. At the integration layer, the same principle applies through a standardized protocol now becoming a common way for agents to reach external tools. An agent should hold its own scoped OAuth token to invoke a given tool, while the downstream credential for the actual backend service stays on the server side, retrieved by the MCP server or application backend rather than handed to the agent directly. If the agent is compromised under that arrangement, an attacker recovers the agent's own scoped session token, not the underlying service credential that token was standing in for.

WorkOS's 2026 guide frames the full credential stack as three layers that have to work together rather than in isolation: encrypted storage at rest, built with per-tenant encryption keys so that a compromised key for one customer cannot expose another's data; OAuth connection management, which fetches scoped tokens at runtime on behalf of the specific user who initiated the action; and session-level access control, which ties every action back to the user who actually authorized it. Each layer closes a gap the other two leave open.

Scoping failures compound as delegation chains grow longer. When Agent A delegates work to Agent B, which in turn calls Agent C's MCP server, credentials can pass through several hops before the task completes, and a compromise at Agent C's server can expose credentials belonging to the entire chain behind it. Scoping at every hop, rather than only at the first one, limits what any single compromised node in that chain can actually reach. That problem deepens further once multiple agents share infrastructure, a question taken up in more detail below.

A related anti-pattern is the shared service account, where several agents authenticate through one identity. Individual agent behavior becomes unattributable under that setup, and revoking access for one misbehaving agent means revoking it for every other agent riding on the same account. Scopehold's 2026 guide recommends the opposite: a durable, individual agent identity for each recurring runtime, with only the specific secrets that runtime's task requires assigned to it.

Diagram: Three Credential Layers That Must Work Together. Visualizes: Show a stacked three-layer diagram representing the full credential architecture WorkOS's 2026 guide describes as needing to work together rather than in isolation.

Why Long-Lived Credentials Are Structurally Incompatible with Agents

Even a credential that is injected at the right time and scoped to the right task becomes a liability if it lives too long. A long-lived credential is a deferred breach: the average detection window for credential compromise runs to months, a pace built for human-speed incident response, not for agents that can act and exfiltrate through a permitted channel within seconds of a successful injection.

Short leases close that gap by design. SPIFFE/SPIRE certificates, already scoped to individual workload identities, carry lifetimes of minutes to hours rather than the weeks or months typical of a traditional service account key, and the SPIRE daemon handles rotation on its own, with no renewal code required anywhere in the application. HashiCorp Vault's dynamic secrets apply the same logic to credentials rather than certificates: a configurable TTL governs the life of each one, and the credential is automatically revoked the moment that TTL expires, with no manual cleanup step left for an operator to forget.

Tooling built specifically for this problem is starting to appear in the open-source ecosystem, including projects like agent-secrets (joelhooks/agent-secrets), which uses age-based encryption to manage exactly this kind of short-lived, per-task secret material. The common thread across all of these mechanisms, from Vault to SPIRE to purpose-built agent tooling, is that none of them trust a credential to outlive the task it was issued for. That is the standard an agent's credentials have to meet: scoped to the task, injected at the moment of need, and expired before anyone, human or machine, has a reason to come looking for them.

Sources

  1. GitHub - joelhooks/agent-secrets: 🛡️ Portable credential management for AI agents — Age encryption, session leases, killswitch
  2. How to manage API keys, tokens, and secrets for AI agents — WorkOS

More in Persistent Agents