Browser Session Persistence in Cloud AI Agents

Browser sessions need VM isolation, not just containers.

Staff Writer · · 10 min read
Cover illustration for “Browser Session Persistence in Cloud AI Agents”
Agent Architecture · October 10, 2026 · 10 min read · 2,219 words

Browser session persistence in a cloud AI agent is not one property but a stack of them, and each layer can fail on its own terms. At the shallowest level sits cookie and authentication persistence: the agent returns to a site already logged in, without repeating a credential exchange it already completed once. One step deeper is browser profile persistence, where cookies, local storage, saved credentials, and user preferences sit together as a named profile that is expected to survive between runs. Deeper still is process-level persistence, where the browser itself never shuts down between tasks: the page stays rendered, the JavaScript heap stays populated, any open realtime connection stays open, and the agent resumes work from where it left off. Below all of that is filesystem persistence, covering downloaded files, cached resources, installed extensions, and working directories that need to still be on disk the next time the agent runs; machine-level persistence through snapshot-restore captures the full state of a virtual machine, its memory, its disk, its running processes, and restores it later so the agent wakes into an environment that is already running.

These layers depend on each other in a strict order, and the dependency runs downward. Filesystem persistence, in turn, cannot survive if the container running the agent is stateless by design and discards its writable layer on every restart. None of these guarantees are properties of the browser or the agent's code. They are properties of whatever the agent is actually running on, and that is the question the rest of this piece works through.

Stateless and container-based environments break session persistence at the infrastructure level

Most failures of browser session persistence trace back to the compute environment, not to anything wrong with how the session itself is managed. Serverless functions and standard containers are built around a specific assumption: that the process will end, that its filesystem will be thrown away, and that the next invocation will start from a clean slate. That assumption is a feature for a lot of workloads. It is the opposite of what a browser agent needs if it is expected to pick up a session where it left off.

Hardening a container, dropping Linux capabilities, applying seccomp profiles, mounting a read-only root filesystem, makes the container harder to break out of, and all of that is worth doing. None of it changes the fact that the container still shares a kernel with its host and with every other container on that host. That shared kernel is the limit. If you don't fully trust the code, or if autonomous agents are acting on live web pages with real credentials, you need a kernel-level boundary as the structural requirement, not an optional upgrade. A browser agent's filesystem and memory need to sit behind a boundary that belongs to the agent alone and to no one else running on that machine.

The browser makes this problem larger. If you hold a browser session durably, across restarts and across untrusted content, you need a compute boundary strong enough to hold real state without exposing that state to anyone else on the machine. That is the requirement a microVM is built to satisfy.

Firecracker microVMs provide the isolation boundary that browser session persistence requires

Firecracker was built at Amazon Web Services to run AWS Lambda and AWS Fargate, workloads that meant running enormous numbers of unrelated customers' code on shared physical hardware, safely, with a fast boot time and a small memory footprint per instance. The design target was strangers: code written by people Amazon had never met, running next to other strangers' code, on machines Amazon operated. That is close to the exact threat model a cloud browser agent faces. The same primitive applies here.

Each Firecracker process runs exactly one microVM, and each of those microVMs gets its own guest kernel, not a slice of a kernel shared with its neighbors. That one design choice removes the shared-kernel escape surface that limits container isolation. The list of things a browser agent actually needs isolated, sessions, the filesystem and its downloads, stored credentials and cookies, network access, the actions the agent takes, screenshots and recordings, logs, and the reset path between tasks, maps cleanly onto a VM boundary instead of needing a pile of container-specific hardening bolted on after the fact.

None of this works without KVM. Production Firecracker deployments need KVM access, either on bare metal with hardware virtualization enabled or on a cloud VM where nested virtualization has been turned on. Firecracker's virtualization path uses a simplified device model that does not support direct hardware device passthrough. A stock Firecracker microVM has no path to a GPU. So this primitive rules out GPU-dependent workloads, a constraint that approaches based on kernel interception do not share, and if you are picking infrastructure for a browser agent that needs GPU access, you need to account for it directly.

The payoff for credentials is concrete. A login can happen once, inside a controlled build step, and the snapshot taken afterward holds a session token that produced it. If a runtime instance running from that snapshot is ever compromised, what leaks is a token that can be revoked, not a password that can be reused anywhere else that password was typed.

Snapshot-Restore and Browser Session Continuity

Snapshot-restore is what makes microVM-based browser sessions affordable and operable at any real scale, but it is not the same thing as a session staying alive, and treating the two as interchangeable produces failures that are genuinely hard to trace back to their cause. A snapshot captures the full memory state of the VM, the state of its block device, every running process including the browser itself, a JavaScript heap that is already populated, and any network connections that were open at the moment of capture. Restoring from that snapshot means resuming a machine that was already running, not rebuilding one from a cold start.

Operationally, that distinction is the whole point. A browser agent waking from a snapshot is already logged in. It is already on the correct page, with the correct DOM already in place, and it does not step back through the setup sequence that got it there the first time. A VM create, in this model, is a restore of a machine that already finished booting once, making sub-second wake latency achievable.

The limitation that catches developers off guard is session expiry. Sessions still expire regardless of what the snapshot preserves. The correct response is to rebuild snapshots on a defined schedule and to build automation that watches for a redirect to a login page and falls back to a full, real login instead of failing in a way that looks successful but isn't.

A second, separate limitation concerns what happens between snapshots, not after them. Snapshot-restore by itself does nothing to protect you if state changes silently in the gap between the last snapshot and an unexpected stop. If the page state moves during that window, whatever changed is gone, because the only state that exists is whatever the last snapshot recorded. Finer-grained autosave patterns close part of that gap: periodic save points taken while the machine is running, on a schedule short enough to matter, so an unexpected stop restores from a recent checkpoint instead of falling all the way back to the original snapshot. These are design decisions with real tradeoffs, and the next section is about choosing among them deliberately.

Choosing between a persistent-alive VM and a hibernated VM for browser agent workloads

The right lifecycle strategy for a browser agent's VM depends on one question: is the workload a session, something continuous that needs to stay live, or a shot, something that fires occasionally and can tolerate a wake-up delay measured in a fraction of a second. Get this wrong in either direction and it costs something real: either unnecessary compute spend or a session that quietly breaks continuity it was supposed to preserve.

A persistent-alive VM stays running at all times. The cost is that compute gets billed continuously, whether the agent is doing anything or sitting idle. During idle time, a hibernated VM costs only storage, not compute, and the wake itself, because it restores from a snapshot rather than cold-booting, still lands in well under a second.

The decision criterion follows directly from the workload. A persistent-alive VM makes sense for an agent genuinely mid-session: one iterating on a working tree of files, one holding open a live WebSocket or another realtime connection, or a REPL whose in-memory state cannot be reconstructed just by reading its prior output. That shift matters most at fleet scale. Hibernation is what makes giving every user their own VM affordable at that scale in a way an always-on fleet cannot match. A middle path also exists for workloads that do not fit cleanly into either category: a long-lived VM with periodic internal save points, available instantly because it never fully stops, but still able to recover from an unexpected interruption without losing the session built up since the last snapshot.

A browser agent infrastructure stack built for serious persistence requirements

Putting the preceding arguments together yields a specific stack, one where a gap at any single layer undoes the guarantees built at every layer above it. Compute sits on bare-metal hosts with real KVM access, running Firecracker microVMs, not shared cloud VMs that cannot nest a second layer of real isolation on top of themselves, and not containers that share a kernel with everything else on the box. Each VM carries its own persistent disk, a block device that survives both sleep and redeploy, holding the browser profile, downloaded files, installed packages, and stored credentials on disk rather than trusting any of it to memory that will disappear.

A snapshot lifecycle governs all of this on a schedule: a defined cadence decides when VM state gets captured, the wake path always restores from the most recent snapshot instead of booting cold, and session token validity gets checked the moment the VM wakes, with a real re-login path standing by for whatever token turns out to be expired. None of this should be visible to the developer building on top of it. Infrastructure overhead at that layer should be zero from where the developer sits.

Security implications of persistence that ephemeral agents never face

Persistence changes the threat model, and that fact deserves more weight than it usually gets. An ephemeral agent that dies and restarts clean cannot carry corrupted state forward, because there is no forward for it to carry anything into. A persistent agent can carry corrupted state forward, and that single difference opens a threat class ephemeral agents are structurally incapable of having.

One concrete version of this involves a persistent personal agent that keeps identity, memory, and tool access alive across sessions, creating a background execution surface, often running on a heartbeat rather than waiting for direct user input, where untrusted content encountered during one run can poison the agent's memory and silently affect every run that comes after it. That threat is tied specifically to persistent agents running background execution, not to persistence in the abstract, and it has no equivalent in an agent that gets torn down after every task.

The same microVM isolation that makes reliable persistence possible also contains these threats, answering both problems at once. The practical discipline that follows is intentionality. If a browser profile, a filesystem, or a process is going to survive across runs, there should be a documented reason it persists and a defined way to reset it, because persistence earns its place as a deliberate choice, not as a default nobody examined.

What developers should verify before trusting persistent browser sessions

Before you hand a cloud environment any browser session state that actually matters, you need answers to a short list of infrastructure questions, not product questions, because those answers decide whether the persistence guarantees described throughout this piece are real or only advertised.

Is the compute primitive an actual VM with its own guest kernel, or a container sharing a kernel with its host and its neighbors? That answer decides whether snapshot-restore can be trusted and whether isolation between tenants is a structural fact or a policy someone hopes holds.

Does the VM run on bare metal with KVM access, or on a cloud VM with nested virtualization genuinely enabled? Without that access, Firecracker's isolation guarantees do not hold in production, regardless of what the surrounding product documentation claims.

Is there a defined snapshot schedule, and does the wake path check session token validity before handing control back to the agent, with a real fallback to a full login when that token has expired? Is each agent or each user given its own VM boundary, rather than sharing a container with other agents or other users? And does the disk holding the browser profile, the credentials, and the downloaded files survive a sleep cycle and a redeploy, or does it reset along with everything else?

These questions are the direct, practical test of whether a browser agent running in the cloud will wake up next time already logged in, on the right page, with its work intact, or whether it will wake up to a blank profile with no memory of what it was doing.

Sources

  1. Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution

More in Agent Architecture