Agent Fleet Management Across Thousands of Tenants
Infrastructure, not orchestration, separates agents that survive in production.

Managing a fleet of thousands of tenant agents, one per customer, running autonomously for weeks at a stretch, is not primarily a question of which orchestration framework sits on top. The failure mode at that scale is architectural: what each agent runs inside, what state it carries between turns, and what it costs while it sits idle.
Managing a fleet of tenant agents is an infrastructure problem, not an orchestration problem
The 2026 field of agentic engineering has converged on a specific distinction. A prompt that performs well once, on stage, in front of an audience, is a different achievement from a system that runs unattended for weeks against real tenant traffic. The gap between those two things is rarely the model. The gap lies in the structural decisions: how memory is scoped, how execution is isolated, how the system behaves when nobody is watching it.
Pattern choice, LLM routing, and the long-running debate between ReAct-style loops and explicit planning are well understood at this point, and they are increasingly commodity. What is not commodity, and what separates a fleet that survives contact with thousands of real tenants from one that doesn't, is what each agent runs inside, how much state it's allowed to carry, how fast it wakes from idle, and what it costs the business while it's doing nothing.
If you run one agent per end-user or tenant, rather than one shared agent for everybody, you take on three infrastructure requirements. Each tenant's state has to be isolated from every other tenant's. Each agent has to wake in well under a second, because nobody tolerates a cold start on a conversational system. And the idle-cost structure has to hold up as the fleet grows from dozens of tenants to thousands, or the business funding the fleet erodes its own margin one customer at a time. The rest of this piece takes those three requirements in order.
What stateless agent architectures do to state at scale
An agent, stripped to its components, consists of a perception module, a memory system split between short-term and long-term, a planning engine, an action executor, and some layer of tool integration. In a fleet where each tenant gets its own agent, every one of those components has to be scoped per tenant. Memory in particular is a first-class structural element, not optional scaffolding sitting on top of the "real" agent, and what happens to it under a stateless architecture is where fleet-scale problems start.
Stateless infrastructure re-bootstraps context on every single invocation. Episodic memory (the record of what happened in past sessions), semantic memory (general knowledge the agent has accumulated), and procedural memory (skills the agent has learned) all have to survive across sessions if they are to be useful, but a stateless architecture cannot hold them there. Every invocation starts cold, so it has to pay the cost of reconstructing context that a persistent agent would already have on hand.
System prompts and re-stuffed context consume a substantial share of all LLM input tokens in the token math of large deployments, leaving less for the actual substance of the task. Multiply that waste by a fleet of thousands of tenants: it becomes a structural tax on every interaction the fleet handles.
The obvious developer instinct is to solve this by stuffing more into the context window: carry the whole history forward, every time, and let the model sort it out. That works fine in a demo with one user and a short session. Vector-indexed episodic and semantic stores are the usual structured alternative, and they bring real trade-offs of their own: retrieval latency, and relevance failures occur when the index surfaces the wrong memory at the wrong time. Those are solvable engineering problems. Stateless-by-architecture is not solvable by better prompting or clever retrieval; it requires changing the infrastructure layer the agent runs on. You make the ephemeral-or-persistent call on an agent's state once, but that single choice forecloses or enables everything downstream, including the user experience and the unit economics of the fleet.
Each tenant agent needs its own isolated virtual machine
Persistent, per-tenant state only pays off if you isolate the environment holding that state from every other tenant's, and that need comes from two pressures that reinforce each other. The second is agent autonomy. If an agent installs packages, runs a browser, or modifies its own working environment, it needs a real kernel underneath it, not a namespace carved out of somebody else's.
The agent design pattern literature from CSIRO names correlated failures and accountability as structural risks in multi-agent systems: this is the kind of risk that emerges when agents share enough infrastructure that one agent's failure can propagate into another's session. Isolation at the virtual machine level is the direct architectural answer to both risks at once: a security boundary and a blast-radius boundary, enforced by the same mechanism.
Containers solve density well, but they were never built to solve isolation at the level a multi-tenant agent fleet needs. A traditional virtual machine gives you real isolation, but it is too heavy and too slow to start for a fleet that needs to spin up and tear down sessions constantly. A container starts fast, but because it shares a kernel with its host, it is too leaky to safely pack untrusted tenant code at fleet density. Firecracker, the open-source virtualization technology built for secure multi-tenant container and function-based services, resolves that tension directly: it delivers hardware-level isolation through a real virtual machine boundary while matching the startup speed and resource efficiency that container-based systems are known for.
The numbers make the case concrete. Firecracker boots in well under a second and carries minimal memory overhead per microVM, and a single server can sustain up to 150 new microVMs every second. With that throughput, thousands of tenant agents can run on a single physical server, each in its own isolated virtual machine, and each can boot or restore fast enough that a tenant never notices the handoff. Amazon Bedrock AgentCore Runtime builds directly on this: it uses Firecracker to run AI agent tools in isolated environments, and the architectural contract is explicit, one session maps to one microVM. Lighter-weight microVM tooling has a role to play for fast, disposable, single-shot execution, but it is a complementary tool for short-lived tasks rather than a substitute for the kind of long-lived, per-tenant isolation a persistent fleet agent needs.
The payoff of per-tenant VM isolation goes beyond security. It changes what the agent is allowed to do. An agent confined to a shared container runtime has to be careful; it cannot install arbitrary packages or risk breaking the environment other tenants depend on. An agent given its own microVM can install packages, break its own environment, and rebuild it, repeatedly, without touching anyone else's session. That's a different and more capable primitive than a shared runtime can offer, available once isolation is enforced at the hardware level.
How snapshot-restore scheduling makes per-tenant isolation economically viable
Giving every tenant a dedicated microVM raises an immediate economic question: how does a fleet of thousands of always-available virtual machines avoid becoming a fleet of thousands of always-billed virtual machines? Snapshot-restore scheduling is the mechanism that turns per-tenant isolation from an engineering ideal into a viable business model.
Agent workloads are bursty by nature. General-purpose cloud compute already runs cold even without this problem: average Kubernetes CPU utilization is in the single digits, and agent sandboxes push that utilization lower still. Most of a fleet's paid-for compute capacity spends its life doing nothing.
Snapshot-restore scheduling resolves the mismatch directly. Boot the microVM once and snapshot the running machine, then restore that exact snapshot on every subsequent wake. The restored agent comes back in well under a second, fast enough to feel like a warm pool was sitting there the whole time, but with no fleet of idle machines actually being paid for in the meantime. An idle tenant agent under this model costs storage, not compute. That's why you can afford to isolate every tenant in its own VM at thousands of tenants rather than dozens.
Snapshot-restore also unlocks a capability that stateless architectures cannot express at all: copy-on-write forking. You can branch a running VM's memory and disk cheaply, and that maps directly onto a pattern agent builders care about: try five candidate fixes from the same starting state and keep only the one that passes. No stateless or container-based architecture can offer that kind of cheap, exact branching, because there is no persistent state to branch from.
A Kubernetes-native path exists for teams who want this model without leaving the Kubernetes ecosystem: DoiT's Agent Substrate multiplexes thousands of stateful agents onto a shared worker pool, suspends the idle ones, snapshots their RAM and local files, and restores that state onto an available sandbox in under a second. That's a genuinely useful pattern, and it comes with a real limitation that shouldn't be waved away: standard tag-based FinOps tooling assumes a stable mapping between a Pod and the team or customer it serves, and once agents are being suspended and resumed onto whichever worker happens to be free, that mapping breaks. Cost attribution, in other words, has to be solved separately from scheduling, not assumed to come along for free.
A 2026 survey of agent sandbox providers lays out the break-even math directly: slot price across providers has converged near a common market rate, snapshotting is the one genuinely free efficiency gain available in this space, and most fleets operate well below the utilization threshold where renting infrastructure stops being cheaper than owning it. That threshold is the number every team building a fleet product needs to know before it commits to a pricing model, which is exactly where the economics go next.
Where fleet unit margins break
Infrastructure decisions made in the previous three sections translate into a single number that governs whether a fleet product survives its own growth: per-tenant infrastructure cost as a share of that tenant's annual recurring revenue. In a healthy fleet, that figure stays in a low single-digit band. Above that band, the product is effectively funding the agent feature out of its own margin. Past a ceiling of 20%, the contract has stopped making sense and needs to be revisited at renewal, because the agent is now costing the business more, proportionally, than the business is charging for it.
Three pricing paradigms are hardening at once across the industry, and each one distributes the economic risk of idle compute differently between the builder and the infrastructure underneath. Pay-per-call pricing is easy to track and easy to explain to a customer, but agents are not economical actors: they retry, they re-query, they loop on a problem longer than a human would, and idle compute keeps running in the background the whole time regardless of how many calls get billed. Outcome-based pricing shifts risk the other direction: you tie payment to results rather than to usage, which protects the customer, but then idle compute at the infrastructure layer must already be cheap, or every unresolved outcome is pure loss.
What falls out of all three paradigms is the same conclusion from different directions: teams that solve idle-cost at the infrastructure layer, through snapshot-restore scheduling and per-storage billing rather than per-compute billing, can undercut any pricing model still built around the assumption that idle compute has to be paid for. That advantage compounds as the fleet scales, because every tenant added to a fleet priced this way adds marginal storage cost rather than marginal idle-compute cost, and the two scale at very different rates.
The design patterns that only work if the infrastructure underneath is persistent
The infrastructure argument made above is not only about cost and security. It determines which agent design patterns are available to build with. The field's taxonomy of production patterns, prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer, and fully autonomous agents, runs along a spectrum from code-controlled workflow at one end to model-controlled agent at the other. The further a pattern sits toward the autonomous end of that spectrum, the more it depends on state surviving across turns, and the more exposed it is to whatever the underlying infrastructure does or doesn't persist.
The Circuit Breaker pattern is a clear example. It monitors agent health, and once a component's error rate crosses a defined threshold, it isolates that component on its own and redirects traffic to backup agents or a degraded mode until recovery. A stateless agent has no error rate to accumulate in the first place, because every invocation starts from zero, so the pattern simply has nothing to monitor.
Without that underlying mechanism, checkpoint recovery stays a diagram on a whiteboard.
Copy-on-write forking, introduced in the economics section as a cost-saving mechanism, turns out to also be a design-pattern enabler. It maps onto the evaluator-optimizer loop, where an agent generates a candidate, critiques it, and revises it within a bounded loop, and onto parallelized speculative execution, where several approaches get tried from the same known-good starting state and only the winner gets merged back. That branching capability is native to VM-level snapshot-restore and unavailable to stateless or container-based environments, which have no persistent state at the right granularity to branch from.
And the correlated-failure and accountability risks named in the CSIRO pattern catalogue for multi-agent systems find their architectural answer in the same place all four of these patterns do: isolation enforced at the VM level, so that one tenant's agent failure has no path into a neighbor's session.

