Per-User Agent Architecture in Multi-Tenant SaaS
Isolating each user's agent in its own microVM prevents resource exhaustion and data leaks.

A production incident makes the failure mode easy to picture. An AI SaaS platform serves hundreds of customers, and each customer's agent executes code against that customer's own data, with every agent drawing from a shared pool of containers. One tenant's agent hits an infinite loop, starves the shared CPU pool, and every co-located tenant sees degraded performance, with interference alone producing 5 to 50 percent performance degradation across co-located workloads according to Blaxel Engineering's analysis of these systems. The postmortem traces back to a shared execution environment that was never built to hold one tenant's workload apart from another's.
Traditional SaaS multi-tenancy never had to solve this problem, because it isolates tenants at the application layer, with separate database rows, separate API keys, and role-based access controls doing the work. That model holds up because the vendor controls every line of code running on its infrastructure. Agentic workloads break that premise at the root: the agent generates and executes code at runtime from a user's prompt, so code the platform never reviewed runs inside the same shared execution environment as every other tenant's workload. OWASP's guidance for LLM applications treats the model as any other user and validates everything it produces, because code the platform never reviewed is untrusted by definition. Resource exhaustion is one symptom of the same underlying weakness. A filesystem or memory access flaw that goes unnoticed at the process level lets one tenant read another's data, and the November 2025 runc breakouts, cited by CNCF analysis, show that container escape CVEs recur on a regular basis and "pose a critical risk" in exactly the multi-tenant environments where users define their own containers or run unvetted images.
Isolation Boundaries: Containers, gVisor, and MicroVMs
Three technologies mark out the range of isolation strength available to a platform building agent infrastructure today. Containers share the host kernel with every other container on the node. gVisor sits a step further out, intercepting syscalls in userspace rather than letting a workload talk to the kernel directly. Firecracker microVMs go further still, giving each workload its own dedicated guest kernel enforced by hardware virtualization.
Kubernetes' own documentation states that containers offer "a weaker isolation boundary" than hardware-based VMs, and the reason is structural: the boundary is the host kernel, and every container on a node shares it. That shared kernel is not a hypothetical risk. CVE-2019-5736 and CVE-2024-21626, both found in runc and both carrying high-severity CVSS scores, each opened a window for crossing tenant boundaries, and an agent stumbling into one of these flaws while executing generated code can trigger it as readily as a deliberate attacker would. MicroVMs close that window by moving the isolation boundary below the kernel entirely: a kernel exploit inside one microVM has no path to the host or to a neighboring tenant's workload, because hardware virtualization enforces the separation rather than application code.
The standing objection to this approach has always been cost: a dedicated VM per tenant sounded like an unaffordable amount of latency and memory overhead to carry at SaaS scale. Firecracker undercuts that objection directly. It boots a microVM in 125 milliseconds or less, according to Blaxel Engineering's analysis of multi-tenant AI agent isolation, removing the latency and memory overhead objection to per-tenant VMs. The question that follows is whether a VM boundary, clearly the strongest option on the spectrum, can actually be run per user at the scale a SaaS platform operates at.
What microVM-per-user isolation looks like at SaaS scale
MicroVM-per-session or microVM-per-user isolation is already the production default chosen by several major providers building these runtimes today.
AWS Bedrock AgentCore, published in a blog post dated 2026-05-21, resolves the dedicated-versus-shared tension through session-isolated microVM-based compute: it launches a lightweight microVM for each session, and that session gets its own ephemeral filesystem by default, with persistent storage available only through optional configuration. That design lets an agent read and write session-scoped files and preserve state across a multi-step interaction without exposing one session's working data to another. A separate infrastructure provider launched a product in January 2026 built around persistent, instantly available VM environments meant to keep AI agents running continuously, using Firecracker microVMs paired with a 100 GB persistent NVMe filesystem and checkpoint/restore capability. The move signals that infrastructure vendors are now designing primitives specifically for agent workloads rather than adapting serverless containers built for something else.
Other platforms occupy a different point on the same spectrum. One widely used sandboxing service runs isolated environments on Firecracker microVMs with kernel-level isolation per sandbox and SDKs for Python and TypeScript, but it is ephemeral by design: sandboxes on its free tier live for only a short window, and even on its paid tier they cap out at 24 hours, with no automatic persistence of state between sessions, though pause and resume can preserve filesystem and memory state. That is the right tool for short-lived code execution and the wrong one for an agent that needs to remember what it installed the week before. CreateOS takes a different position again, describing itself as a unified AI execution layer that runs agent workloads in Firecracker and KVM micro-VM sandboxes, with fork and pause-resume support, VPC and S3-backed disks, kernel-level egress control through eBPF, and the option to run entirely inside a customer's own infrastructure boundary.
The economics behind all four of these approaches rest on a density number that makes per-user VM deployment practical rather than exotic: Firecracker's memory overhead allows a large host to support thousands of simultaneous microVMs. The operational security model benefits from a related design choice: Firecracker runs one VMM process per microVM, rather than a single daemon managing every VM on a host, so compromising one VMM does not touch any of the others. Together, the density figure and the per-VMM isolation model are what turn VM-per-user from a theoretical ideal into something a platform can actually run at the scale a SaaS business operates at.
The four architectural decisions that follow once you commit to per-user agent isolation
Committing to a VM-per-user model does not end the design work, it starts a new phase of it. Four concrete decisions follow from that commitment, each one deferred or avoided entirely by shared-runtime designs: how tenant identity flows into the execution boundary, how state persists across sessions, how service components get decomposed for independent scaling, and how to handle the tradeoff between shared and separate data stores.
Tenant identity has to flow explicitly into each isolated execution environment. It cannot be inferred from a shared connection context the way it often is in a traditional multi-tenant application, because there is no shared context left once each tenant's agent runs in its own VM. AWS Bedrock AgentCore handles this with custom HTTP headers that carry tenant-specific metadata, including a tenant identifier, tier, regional preferences, feature flags, and entitlements, alongside the standard authorization tokens. The agent reads these headers at invocation time, which lets it run workflows tuned to that tenant's own business logic without hardcoded routing baked into the agent itself. The cost of getting this wrong is not abstract. Scalekit documented an isolation incident caused by a single GitHub token shared across tenants, resolved only once tenant-scoped agent identities backed by IAM were put in place. That incident is the concrete failure mode that makes explicit identity propagation a hard requirement rather than a nice-to-have optimization.
Persistent state is the second decision, and it has real consequences for what an agent can do across sessions. Each agent needs a filesystem that survives sleep, redeploy, and the scale-to-zero boundary, because without it the agent loses every installed package, credential, and intermediate computation artifact between sessions. AgentCore Runtime can give each session its own persistent filesystem that survives stop and resume cycles, but only once managed session storage, currently in public preview, is explicitly configured; left at its default, every session boots into a clean, ephemeral filesystem.
Service decomposition for scaling is the third decision, and it follows from the fact that not every component in a per-user agent system scales the same way. The agent runtime and the tool executor are stateless, so they scale horizontally, adding more instances as load grows. Config and conversation services are stateful but comparatively low-traffic, so they scale vertically instead, through database optimization rather than horizontal replication. Getting this decomposition wrong pulls a platform in two opposite directions at once: either stateful services end up over-provisioned while execution capacity runs short, or the reverse happens, and neither failure mode is cheap to unwind once traffic patterns are already live.
The fourth decision concerns the data layer itself, and it runs along what the refact.co guide on multi-tenant SaaS architecture describes as the pool, bridge, and silo spectrum. A shared schema with row-level security, the pool model, is the sensible default for most B2B SaaS platforms. A schema-per-tenant arrangement, the bridge model, fits a narrow band of compliance needs. Database-per-tenant, the silo model, gets reserved for the largest accounts, white-label deployments, or workloads under heavy regulation. Whichever model a platform chooses, tenant_id belongs on every business object from day one, because retrofitting it later is painful, and isolation enforced only by application code erodes with every merge that touches the schema. Row-level security in Postgres is a necessary piece of this but not a sufficient one on its own: connection pools reuse sessions, so RLS needs to be scoped per-transaction, through SET LOCAL or an equivalent mechanism, or the wrong tenant's context can bleed through a reused connection.
How retrieval and tool authorization break in multi-tenant agents
None of the compute isolation covered so far secures retrieval or tool execution on its own. A retrieval system that ranks documents by relevance rather than by authorization can surface one tenant's confidential data to another tenant because that document scores highest against the query, even when the requester is not allowed to see it. A VM boundary stops one tenant's code from reaching another tenant's process. It does nothing to stop a retrieval layer from handing over the wrong tenant's documents because nothing in the ranking logic ever checked who was asking.
The Red Hat AI paper "Securing the Agent," published in May 2026 under arXiv:2605.05287, names this gap formally and identifies three further failure modes beyond the mismatch between relevance ranking and authorization: tool-mediated disclosure, where a tool call exposes data its caller was never cleared to see; context accumulation across turns, where information from one tenant's session leaks into the working context of another; and client-side orchestration bypass, where logic running on the client sidesteps authorization checks meant to live on the server. Replicating the entire agentic stack per tenant, with separate vector stores, dedicated inference endpoints, and isolated tool configurations for each one, is the naive response to this gap, and it fails on cost: infrastructure spend scales linearly with tenant count instead of with actual usage, and the operational fragmentation that results only compounds as the tenant list grows.
The paper's proposed alternative is a layered architecture combining policy-aware ingestion, retrieval-time attribute-based access control (ABAC) gating, and shared inference, all enforced through server-side agentic orchestration. Security-critical operations, meaning tool execution authorization, state isolation, and policy enforcement, stay centralized on the server, while client-side frameworks keep control over agent composition and over the latency-sensitive operations that benefit from staying close to the user. Client-side orchestration still wins on throughput at high concurrency for latency-sensitive workloads, but only where every tenant involved is already trusted. It cannot enforce an authorization boundary between tenants who are not supposed to trust each other, and that enforcement has to live on the server.
MCP as the emerging standard for tenant-aware tool integration in agent SaaS
Tool integration for AI agents is converging fast on a single standard, and that convergence means SaaS builders can no longer treat it as a one-off integration problem solved differently for every tool. A leading tool-integration protocol crossed tens of millions of monthly SDK downloads by March 2026, moving from an internal experiment at a major AI lab to a Linux Foundation project in roughly 13 months, a pace of standardization that marks it as infrastructure rather than an experiment still finding its footing. By April 2026, ten major AI agents supported custom remote MCP servers with native OAuth 2.1: Claude and Claude Desktop, Claude Code, ChatGPT, VS Code with GitHub Copilot, Zed, Kiro from AWS, Amazon Q Developer CLI, OpenCode, Docker MCP Toolkit, and Cursor.
The transport choice behind an MCP deployment carries real architectural weight. Stdio remains best for local process integration and CLI agents, while Streamable HTTP is the default for modern cloud-hosted MCP. That shift to HTTP means MCP servers now have to handle a set of concerns that used to sit with the client: authentication, rate limiting, multi-tenancy, and authorization all move onto the server once the transport does. An MCP server running in a SaaS context becomes a new enforcement surface for the same tenant isolation requirements built out across the rest of this architecture, carrying the identity-routing and authorization-gating work described above. AWS Bedrock AgentCore already treats this as a first-class concern rather than an afterthought, providing constructs for hosting MCP servers directly within its multi-tenant managed runtime. The tool integration layer and the compute isolation layer are converging on the same set of requirements, because a platform that gets tenant identity, state persistence, and authorization right in its execution environment has to carry those same guarantees into every tool call an agent makes on a tenant's behalf.
Sources
- Building multi-tenant agents with Amazon Bedrock AgentCore
- Multi-tenant AI agent isolation for SaaS platforms
- Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use
- Multi-Tenant SaaS Architecture: A Practical Guide
- Access Control for Multi-Tenant AI Agents: Identity & Isolation
