Per-User Agent Architecture for SaaS Products

Agents need isolated compute, memory, and tools—not shared infrastructure.

Reporter · · 10 min read
Cover illustration for “Per-User Agent Architecture for SaaS Products”
Agent Architecture · October 2, 2026 · 10 min read · 2,267 words

SaaS builders now face an architectural decision they cannot push to next year's roadmap: whether each user on the platform gets a dedicated agent, or whether the whole user base shares one. The industry's center of gravity has already moved away from copilots, tools built to suggest, draft, and assist, toward agents that research, act, and iterate without a human approving each step. That shift changes what the software does in the gaps between sessions, when no one is watching the screen, and that alone makes the decision structural rather than cosmetic. Gartner's forecast has the share of enterprise applications embedding task-specific agents moving from a small fraction today to a substantial minority by the end of 2026, which puts this decision inside the current release cycle, not a planning exercise scheduled for 2027. Satya Nadella's framing, as reported by Glean, holds that business applications as they exist today will collapse in the agent era, with AI systems updating multiple databases directly and business logic moving into the AI tier itself. If that holds, a shared, stateless AI layer bolted onto an existing application isn't a lighter-weight option, it's a layer that cannot do the job the agent is now expected to do, because the agent is becoming the application rather than a feature sitting on top of it.

What a shared, stateless AI feature cannot do

A wrapped LLM endpoint forgets the previous conversation the moment it ends, carries no tools of its own, and leaves nothing behind that resembles an audit trail. Vendasta's 2026 production guide draws the line precisely: the gap between a chatbot wrapped around an LLM call and an agent that can actually book an appointment, update a CRM, escalate a complaint, and explain afterward why it did so. That gap appears in production within days, through three distinct failure modes.

The first is memory bleed. Without per-tenant isolation, one user's session context leaks into another's, a risk built into any shared memory store that relies on a default or global partition key. The anti-pattern that produces this failure in practice is setting MEMORY_TENANT_ID=default at deploy time instead of passing the tenant ID through request headers and partitioning the memory store per tenant at the storage layer. The second is quota starvation: in shared infrastructure, a single noisy tenant can exhaust the LLM API quota for everyone else on the platform. The orchestrator's execution pipeline needs per-tenant token budget gates built in from the start, not added after the first outage. The third is context collapse. An agent that cannot persist its working state between sessions cannot carry a multi-step workflow across turns with a user, so every conversation restarts from zero regardless of how much work came before it.

All three share a harder cause: standard multi-tenant containers cannot safely contain AI agents that execute arbitrary code. Once an agent can write its own Python scripts, install packages, and manipulate file descriptors, the shared kernel surface area of a standard Docker container turns from a minor inconvenience into a direct security liability. Teams that treat this as a configuration problem, solvable with better namespacing inside the same container runtime, are solving the wrong layer of the stack.

The six-layer stack a per-user agent requires

Diagram: The Six Layers a Per-User Agent Requires. Visualizes: Visualize the six infrastructure layers that production-grade per-user agents depend on, as enumerated in Vendasta's 2026 guide: (1) Compute & Sandboxing, (2) Memory, (3) Tools &…

Production-grade per-user agents depend on six distinct infrastructure layers, drawn from Vendasta's 2026 infrastructure guide for SaaS AI agents, and skipping any one of them is usually where a rollout stalls. Compute and sandboxing runs agent-generated code in isolation; cold-start latency, escape paths, and regional residency are the details teams building this themselves consistently underestimate. Memory stores the working, episodic, semantic, and procedural context an agent needs to recall a given user, their account, and what it did for them last time; per-tenant isolation, retention rules, and the cost of embeddings at scale are where the hidden complexity lives. Tools and actions let the agent call APIs, browse, send messages, book meetings, and update systems of record, and the failure modes here are auth-on-behalf-of handling, rate limits, idempotency, and version drift as those downstream APIs change. Model routing selects, routes, and falls back across language models depending on the task at hand and the cost ceiling set for it. Orchestration sequences single- and multi-agent workflows, handles retries and branching, and manages long-running tasks, with state recovery after a crash and human-in-the-loop handoffs the two places teams get it wrong. Observability and governance trace every decision an agent makes, enforce guardrails, and produce an audit log, with jurisdictional logging and accuracy regression alerts the pieces most often left out of a first build.

The orchestration pattern that has become canonical in production is the Orchestrator-Worker-Critic triad, and LangGraph remains the most widely deployed framework for it, because its state machine model maps cleanly onto multi-agent workflows and its persistence layer answers the question of what happens when the agent crashes mid-task. Most agent projects that stall do so because the failure rate concentrates in the orchestration and isolation layers, not because the underlying model isn't capable enough. Teams that approach agents as a model-selection exercise, rather than an infrastructure build, tend to hit this wall around month three, when the bills arrive and the edge cases that a demo never surfaced appear in production.

Compute isolation as the non-negotiable foundation of per-user agents

Of the six layers, compute and sandboxing carries the most weight, because once an agent can write and execute arbitrary code, install packages, and manipulate file descriptors, container-level isolation stops being adequate. Micro-VM isolation is the correct primitive for this workload, not a more cautious alternative to it.

Firecracker, the open-source microVM technology, has become the standard for high-security agent sandboxing. It boots a dedicated Linux kernel per sandbox in approximately 125 milliseconds, with a device surface area dramatically smaller than what full virtualization exposes. The argument for this over container-based isolation is decisive once an agent is writing its own scripts: a Docker container's shared kernel becomes a liability the moment that capability exists, and the exposure isn't a theoretical edge case but a direct consequence of what the agent is now able to do. DigitalOcean's move to launch MicroVMs in private preview, a compute layer built on Firecracker and aimed specifically at AI agent infrastructure, sandboxes, and short-lived workloads, with automatic pause-on-idle and resume that brings memory, files, and processes back intact, is a signal that microVM-backed agent compute is becoming a standard offering rather than a specialized niche product.

GPU-bound agent workloads sharpen the distinction further.

What per-user isolation buys a SaaS builder is concrete rather than abstract. Each agent gets its own kernel, so it can install packages, break something, and rebuild without touching any other user's environment. Persistent disk state carries over between sessions, so files, credentials, and intermediate computation artifacts are waiting when the agent wakes rather than needing to be reconstructed. And because isolation is strong at the kernel level, a runaway process, a crash, or a security event in one user's agent stays contained to that one environment instead of spreading across the fleet.

State persistence as the architectural requirement that follows from per-user isolation

Isolation without persistence produces an agent that forgets everything the moment it goes to sleep, undermining the purpose of giving it a dedicated environment. A per-user agent that resets on every session is, functionally, a shared agent with extra infrastructure underneath it. The isolation investment only pays off once the agent's state survives between the moments it's actually running.

A per-user agent needs to maintain four distinct types of memory. Working memory holds the current task context and whatever tool calls are in flight. Episodic memory records what the agent did in prior sessions with this specific user. Semantic memory holds what the agent knows about the user's domain, their preferences, and their account. Procedural memory captures the workflows and tool sequences the agent has learned work well for this particular user over time.

Firecracker's snapshot-restore mechanism is the operational foundation that makes this possible for agents whose context accumulates across dozens of tool calls. Pausing a sandbox saves both its filesystem state and its in-memory state, a paused sandbox can sit indefinitely without further action, and billing stops entirely while it's paused. Whether in-memory continuity actually survives depends on how different platforms implement the restore side of that equation. Snapshot-based approaches that restore into a new sandbox rather than resuming the original one have a known limitation: filesystem snapshots can outlive the 24-hour maximum lifetime, but the in-memory session state does not come back with that restoration.

Morph's Infinibranch approach takes a different route, built for agents that need to branch a single environment into many parallel copies continuing from the same starting state: it snapshots an entire running VM and branches or restores it in under 250 milliseconds, a speed that matters for parallelized agent workflows sharing an initial context. Amazon Bedrock's AgentCore Runtime runs each session in its own isolated filesystem on microVM-based compute, with persistent filesystem state across stop and resume available as an opt-in preview feature rather than the default behavior, letting agents maintain intermediate computation artifacts and carry state across multi-step interactions while reducing the risk of cross-session data leakage and keeping the cost profile manageable. Each of these represents a different trade-off between how much state survives a pause and how fast the agent can come back, and a SaaS builder needs to evaluate that trade-off against what their own agent actually does between sessions.

Lifecycle management: how per-user agents sleep, wake, and stay economically viable

Per-user agents exist continuously as concepts, holding memory and state tied to a specific user, but they cannot run continuously as compute without the cost model collapsing. Most agents are idle most of the time, and that gap between user turns is the single biggest hidden cost in interactive AI applications, where only a fraction of billed time is ever spent on productive computation. The cost model that follows from that observation is straightforward: an agent should pay for storage while idle and pay for compute only while active. That's what makes deploying agents per-user economically viable at scales running into the hundreds or thousands of concurrent agents.

Snapshot-restore scheduling is the mechanism that puts this into practice. The micro-VM pauses, its memory and filesystem state get preserved, billing for compute stops, and when the user returns, the sandbox resumes in tens of milliseconds with its full context intact. Wake latency is a hard product requirement here: if a user perceives a multi-second delay between taking an action and the agent's first response, the product feels broken regardless of what the agent eventually does. Sub-second wake is the bar a per-user agent fleet has to clear.

At fleet scale, the warm-pool pattern addresses the first-request problem directly: a pool of pre-initialized sandboxes sits ready so that no user's first request has to wait through a cold boot. Running this safely for agents that execute arbitrary code requires gVisor-sandboxed Kubernetes pods or an equivalent isolation layer underneath the warm pool rather than a bare container scheduler. Platform choice at this layer has an outsized effect on unit economics. A 2026 cost comparison across infrastructure platforms found meaningful differences for equivalent agent workloads, with managed platforms carrying a significant premium over bring-your-own-compute approaches. For a SaaS product pricing per-user agents into its subscription tiers, that premium either lets margin hold at scale or erodes it as the user base grows.

Three multi-tenancy patterns for agent workloads

The fully siloed pattern gives each tenant its own dedicated compute, memory, and storage stack. Isolation under this pattern is as strong as it gets, since nothing is shared between tenants at any layer. The cost and operational overhead of running separate infrastructure per tenant make this viable only for high-value enterprise contracts, where the contract size justifies the dedicated footprint; it does not scale to a SaaS product serving hundreds or thousands of users on a typical subscription cost model.

The fully shared pattern relies on logical isolation through a tenant_id field inside otherwise shared infrastructure. It's the cheapest option to run and the simplest to operate, so many teams default to it early, which is also where the memory bleed failure mode described earlier actually originates: a single global partition key, the MEMORY_TENANT_ID=default anti-pattern, lets one tenant's data contaminate another's, and that is a production failure, not a hypothetical risk. A single noisy tenant can exhaust the shared LLM API quota for every other tenant on the platform unless the orchestrator enforces per-tenant token budgets. And for any agent capable of executing arbitrary code, this pattern's shared kernel surface area makes it a security liability regardless of how carefully the tenant_id logic is written.

The pattern that holds up at production scale is hybrid namespace isolation. Infrastructure is shared for cost efficiency, the same way it is in the fully shared pattern, but data is separated using real logical boundaries rather than a single flat field: namespaces inside vector databases, workspaces inside file systems, and per-tenant token budgets enforced inside the orchestrator. The tenant ID travels in request headers rather than getting fixed at deploy time, and the memory store partitions per tenant at the storage layer itself. This pattern keeps the economics of shared infrastructure while closing off the isolation failures that make the fully shared pattern unsafe for agents handling real user data and executing real code on a user's behalf. For a SaaS builder deciding how to structure a per-user agent fleet, hybrid namespace isolation is the pattern that matches the actual shape of the problem: shared cost, separated risk.

Diagram: Three Multi-Tenancy Patterns: Cost vs. Isolation. Visualizes: Show the three multi-tenancy patterns for agent workloads on a two-axis spectrum — isolation strength (low to high) versus operational cost/complexity (low to high): Fully…

Sources

  1. AI Agent Infrastructure for SaaS in 2026: Ship AI in Weeks, Not Quarters
  2. Will AI agents replace SaaS? Key insights for 2025
  3. DigitalOcean MicroVMs: the compute foundation for building your AI agent infrastructure (private preview)

More in Agent Architecture