Designing AI Agents

Stateful vs Stateless Agents in Multi-Tenant Products

Stateful agents need isolation by design, or multi-tenant systems leak data and lose state.

Contributing Editor · · 10 min read
Cover illustration for “Stateful vs Stateless Agents in Multi-Tenant Products”
Per-User Agent Patterns · October 7, 2026 · 10 min read · 2,330 words

A developer builds an agent, tests it alone for a week, and it works beautifully: it remembers the last conversation, picks up a task where it left off, handles follow-up questions without re-explaining context. The team adopts it. Within hours, something breaks: one user sees another user's data, a session forgets everything mid-task, costs spike without warning. None of this is a bug in the traditional sense. It's the result of a decision made early and never revisited: whether the agent holds state, and if so, how that state is kept apart for each tenant. Framing this as "does my agent remember things?" undersells what's actually being decided, which is whether each user gets an isolated execution context or quietly shares one with everyone else. A single-user agent can get away with one thread and one context because there's no second party to collide with; that simplicity is a convenience of scale, not evidence the design is sound. Adding a second user breaks the underlying assumptions: single-tenant and multi-tenant agents share a surface vocabulary but demand almost entirely different engineering, from how memory is stored to how requests get routed. The failure modes that follow, cross-tenant context leakage, session state stranded on the wrong server, costs that multiply in ways nobody modeled, rate limits tripped by aggregate load, simply don't exist when there's only one user to serve. Worse, the decision is hard to undo. Retrofitting durable session memory onto a system built stateless typically means rebuilding the orchestration layer.

Where stateless agents belong

A stateless agent treats every request as a self-contained transaction: input goes in, a prompt gets built, the model responds, and nothing is kept afterward. It suits tasks like extraction, classification, and single-turn question answering, where each call has everything it needs and nothing carries forward. Stateless designs scale horizontally behind simple load balancers, skip session synchronization entirely, keep latency overhead low, and close off any cross-tenant leakage surface because there's no persisted data to leak. If a regulated industry can't retain user data between calls, it gains a genuine compliance asset, not just a limitation to work around. The cost appears elsewhere, in what prefix caching gives up. A stateless design that resends the full conversation context with every call forfeits the benefit of prefix caching, the practice of recognizing a stable portion of a prompt and caching it so the model doesn't reprocess it from scratch. That caching cuts time-to-first-token and lowers per-call token cost in systems built to take advantage of it, and statelessness trades that saving away in exchange for architectural simplicity. Stateless models also break down the moment a task requires more than one step: tool outputs need to be stored, checked, and carried into the next step, errors need tracking, and retries need some memory of what already failed. None of that fits inside a model that forgets the instant it answers.

State in stateful, multi-tenant agents

Calling an agent "stateful" means a runtime wraps the underlying language model and manages session data, memory, and tool state on its behalf. The practical default in production systems is a hybrid: the LLM call itself stays stateless, and continuity comes from the layer around it. State belongs to the runtime, not the model, and that distinction matters because it tells you where to look when isolation fails. Three memory tiers have become standard in this kind of runtime. Working memory is the active context for a single call, the message buffer, tool traces, and system prompt, re-derived each time and discarded once the request finishes; it carries no cross-tenant risk because it never persists. Episodic memory is a time-ordered record of past interactions, usually backed by a vector store or a Redis key-value store, scoped to a session or a user. External-memory systems outperform full-context-window baselines by 10 to 30 percentage points on memory-specific benchmarks, per 2026 analysis. Semantic memory holds consolidated facts, preferences, and profile data in a structured store or knowledge graph, and it's both the longest-lived tier and the most sensitive one in a system serving many tenants at once. None of this works if the payloads get bloated. Keeping persisted state compact, separating what's transient from what actually needs to sync, and defining an explicit initial state aren't stylistic choices; you need them if the system is going to hold up under load. The operational loop looks the same across implementations: a request arrives carrying an entity key, user_id, session_id, or workflow_id, the runtime loads the stored context tied to that key, combines it with the new input, calls the model and any tools, computes the updated state, and writes it back. Every step in that loop touches a store that other tenants are also reading from and writing to. The next set of problems begins there.

How stateful multi-tenant agents fail

The most damaging failures in a multi-tenant stateful agent rarely look like the model is behaving badly. They trace back to isolation and routing decisions made long before anyone wrote a prompt. Context leakage sits at the top of the list for sheer damage: memory that isn't scoped per user, session IDs that get shared or reused, and conversation history appended to one global buffer shared across users. In a single-user deployment, none of that registers as a problem, since there's no one else's data to collide with. In a multi-tenant deployment, even a small leak can expose one account's sensitive information to another. A quieter failure follows close behind: localized amnesia, where session history gets stranded on the specific server instance that handled earlier turns. A load balancer sends the next request somewhere else, and the agent appears to forget everything, even though the data still exists on the original instance. It looks like a memory bug. It's a routing failure wearing a memory bug's clothes. Underneath the state store itself, five concrete failure modes recur: stale state left behind by parallel overwrites, partial updates that never complete, race conditions between simultaneous writes, prompt drift as context accumulates inconsistently, and state lost outright across retries. Systems that depend on developers remembering to attach a metadata filter to every query carry more risk than systems built to reject out-of-bounds queries at the query planner itself, because the first depends on human discipline holding under pressure and the second doesn't. Concurrency is where the real difficulty lives: multiple tenants writing state at the same time create consistency hazards that single-user testing never surfaces and that staging environments rarely reproduce faithfully. Cost adds a final layer of pressure. One user sending 20 requests a day is trivial to absorb. Hundreds of users doing the same thing, each triggering multi-step reasoning chains, raise cost totals and rate-limit collisions that stay invisible in projections until they appear on an invoice.

The four architectural layers where isolation must be designed in, not bolted on

Diagram: The Four Layers Where Tenant Isolation Must Hold. Visualizes: Show four sequential architectural layers that must each enforce tenant isolation independently: Routing (session pinning or centralized cache via Redis, entity key on every…

Tenant isolation isn't a single switch to flip. It has to hold at four layers simultaneously, routing, storage, execution, and observability, and a gap at any one of them undermines the other three no matter how well they're built.

At the routing layer, every request needs to carry a tenant-scoped entity key, whether that's user_id, session_id, or workflow_id, and the load balancer sitting in front of the agent needs to either pin sessions to a specific instance or route through a centralized cache. Redis or session pinning fixes localized amnesia directly, because the stored state becomes available to whichever instance ends up handling the next request, instead of living only on the one that handled the last.

At the storage layer, each memory tier needs its own isolation strategy, not one blanket rule. Working memory, scoped to a single request and never persisted, carries no cross-tenant risk by construction. Episodic memory needs a key schema like session:{channel_id}:{user_id}, because that key is the actual isolation boundary, and not just a convention layered on top of it. Semantic memory is the highest-risk tier of the three: structured stores and vector indexes need tenant filtering enforced at the query level, not left to application code that a developer might forget to write correctly on a bad day.

At the execution layer, any task that touches a file system, installs packages, opens a browser session, or runs code needs compute isolated per tenant. Shared containers or shared serverless functions can't guarantee that boundary holds. Per-session or per-user micro-VMs can. AWS Bedrock AgentCore Runtime addresses this with session-isolated, microVM-based compute: it launches a lightweight microVM for each session, gives that session its own isolated file system, and can preserve state across multi-step interactions, with persistence across stop and resume cycles available through opt-in managed session storage currently in preview, which reduces the risk of one session's data bleeding into another's. A separate research effort, DeltaBox, built by Shanghai Jiao Tong University and Huawei (arXiv:2605.22781), achieves checkpoint and rollback at millisecond speed, 14 milliseconds to checkpoint and 5 milliseconds to roll back, by duplicating only the incremental changes between consecutive checkpoints. That makes branching and rollback practical at the granularity of individual agent steps, without the overhead that full-state snapshots would otherwise impose. Maritime applies a related idea to multi-tenant products directly: each agent runs in its own Firecracker micro-VM on bare metal, with a persistent disk, a real kernel, and dependencies preloaded, so the isolation boundary is the VM itself, a filter a query might bypass. Its snapshot-restore scheduler lets a sleeping agent wake in under a second, so you get strong per-user isolation without paying for compute that sits idle around the clock.

At the observability layer, every LLM call, tool call, and state mutation needs to be tied to a persistent trace ID. If trace correlation isn't scoped to each tenant, you can't investigate a leakage incident or answer a compliance audit once the system is live. AWS has formalized three patterns that map onto this layered approach: the silo pattern, which gives each tenant dedicated agent skills at the cost of separate maintenance for each one; the pool pattern, which shares agent skills across tenants; and the bridge pattern, which places common steps like authentication, logging, and error handling into shared skills that then invoke tenant-specific skills at runtime.

Per-user VM isolation economics

If you think strong per-tenant isolation costs too much, that belief rests on a pricing model built around always-on compute. Snapshot-restore scheduling removes that constraint, and it changes what per-user VM isolation costs at scale. Start with a number most planning documents miss: the total cost of running a production AI agent typically runs two to five times higher than the raw model API cost that teams estimate going in, because infrastructure, orchestration, storage, retrieval, observability, and maintenance make up most of the bill. Agentic workflows also burn through far more tokens per completed task than a standard chatbot does, because of orchestration overhead, retries, and context that has to be resent in full on every call under a stateless design. Idle compute adds another layer of waste: an agent spends most of its time waiting for a user to do something, and a pricing model that charges for a running VM through all that waiting wastes most of what gets paid for. Durable execution frameworks address part of this at the orchestration layer: a workflow that's fully suspended while idle consumes no compute at all, so suspend-and-resume patterns cut the cost of that idle time substantially. Maritime's snapshot-restore scheduler takes the same idea down to the infrastructure layer: an idle agent pays for storage, not compute, because its VM gets snapshotted while sleeping and restored in under a second when a request comes in. That shifts the dominant cost from idle VM time to storage, so you can give every user their own persistent agent instead of saving it as a luxury for a few high-value accounts. The same architecture removes the metadata-filter risk discussed earlier outright: there's no shared query planner to misconfigure, because each agent's persistent disk is physically separate from every other tenant's. Stateless designs that dodge this cost by avoiding state entirely pay for it elsewhere, in more round trips per task, higher token spend per resolved request, no prefix-caching benefit, and a degraded experience for any workflow that depends on continuity across turns.

The tooling layer, MCP, CLI, and SDKs, and stateful multi-tenant isolation

The protocol an agent uses to reach external tools is not a separate decision from the isolation architecture surrounding it. MCP, CLI-based tool access, and SDK integrations each impose their own tradeoffs on context budget, credential handling, and audit compliance once multiple tenants are involved. The stakes are concrete: connecting a standard GitHub MCP server dumps roughly 55,000 tokens into context before the agent has done anything useful, per 2026 analysis. In a stateless design, you pay that cost on every single call, because nothing is cached between requests. In a stateful runtime with prefix caching and session-scoped tool state, that overhead gets paid once per session rather than once per turn, which changes the math on which tooling approach is affordable at scale. Credential management follows the same logic: a CLI-based tool invoked inside a shared container has to isolate credentials per tenant through application logic, while a tool invoked inside a per-user micro-VM inherits isolation from the execution boundary itself, since one tenant's process has no path into another tenant's disk or environment variables. Audit requirements compound this further. A tool call made through MCP, CLI, or an SDK needs to be correlated back to the same trace ID and entity key used everywhere else in the system, or the observability layer described earlier breaks down exactly where tool use intersects with user data. None of this is a detail to defer until after the agent works. The tooling layer sits inside the same isolation boundary as memory and execution, and treating it as an afterthought reopens the exact leakage and audit gaps the rest of the architecture was built to close.

Sources

  1. One post tagged with "stateful vs stateless agents"

More in Per-User Agent Patterns