Designing AI Agents

Onboarding Flows That Provision a Live Agent per User

Every new user's agent must be fully provisioned before their first interaction.

Staff Writer · · 10 min read
Cover illustration for “Onboarding Flows That Provision a Live Agent per User”
Per-User Agent Patterns · October 8, 2026 · 10 min read · 2,319 words

Most teams still treat onboarding as a design problem: a welcome email, a profile form, a progress bar that makes the first five minutes feel good. That model holds up fine for software where the "agent" is really a shared inference endpoint with no memory of the person typing into it. But if a product promises each user a persistent agent, one that remembers them, acts for them, and picks up work between sessions, that model falls apart. Research on the shift from conversational chatbots to what's been called "digital colleagues" frames this as a move from answer generation to authorized work delegation: the agent decides what intermediate steps to take, not the user, and that only works if the agent carries its own execution context forward in time. Once a product commits to that model, signup stops being a database write. It becomes a provisioning event: a virtual machine has to come up, a filesystem has to initialize, an identity has to be established, and credentials have to get wired in, all before the user's first real interaction. A team might object that this can be faked with a shared backend and per-user rows in a state table, but that approach collapses under real multi-tenant agent load: concurrent file writes, package installs, live browser sessions, and long-running subprocesses all need real isolation, not a shared process pretending to be several.

Components of a Per-User Agent Environment

A production agent assigned to one user is not a process that runs and exits. It has four parts, and each one has to survive the gap between sessions instead of resetting on every invocation. The first is compute isolation: the agent needs a kernel of its own, so it can install packages, open browser sessions, and run code the model wrote without touching anyone else's environment, something a shared container cannot guarantee. The second is a persistent disk: files, cached state, installed tools, and credentials need to still be there when the agent wakes back up, because an environment that resets to a blank filesystem forces the agent to relearn its own setup every time it's called. The third is a standing identity. By 2026, agent authentication treats the agent as its own identity class, not a human account or a service account, because its permitted scope changes with every invocation, and it has to carry a delegation chain showing which user authorized which action. The fourth is a memory store: before a conversation starts, the agent looks up a persistent record keyed to the user, pulls in past work, preferences, and credentials already in use, and loads that into its working context; once the session ends, it writes a summary back for next time. The "Workspace and Skill" framing used in the agent research literature treats this persistent workspace as the substrate that turns episodic tool calls into something that resembles a colleague's ongoing work: state that doesn't disappear, procedures that get reused, tasks that actually close out, and experience that carries forward. None of this happens automatically. So how do you stand up compute, disk, identity, and memory in the two or three seconds a user spends staring at a signup confirmation screen?

The latency constraint: why provisioning must complete before the user's first interaction

If a newly signed-up user's agent isn't ready the instant they open the product, the product has quietly reverted to the stateless chatbot experience it was supposed to replace. That sets a hard limit: the whole provisioning pipeline has to finish inside the signup flow's existing latency budget, measured in seconds, not minutes. Booting a fresh virtual machine from scratch for every new signup doesn't fit that budget; a cold boot adds delay a signup confirmation screen was never designed to absorb. The fix is snapshot-restore. Using Firecracker's snapshot-restore API, a microVM gets booted once to a fully ready state, with memory and block device state snapshotted to local NVMe storage; every new user's VM then restores from that snapshot rather than booting fresh, landing in the exact state, filesystem included, that the original snapshot captured, with sandbox creation from snapshot completing in under 30 milliseconds. Firecracker's own cold boot runs about 125 milliseconds and carries roughly 5 MB of memory overhead, so even a full cold restore stays comfortably inside a one-second budget. The "golden image" snapshot becomes the artifact the whole pipeline restores from: when the base agent environment changes, a new package or an updated tool, the fix is to re-snapshot that base image once, covering every existing user's VM without re-provisioning each one individually. The remaining pieces add almost nothing to the clock. Creating a memory store record for a new user is a single write, an empty record keyed to the user's ID, and issuing a short-lived delegation token is a synchronous API call that runs in parallel with the VM restore. Both finish in milliseconds, not seconds.

Designing the signup-to-agent pipeline: the deterministic sequence every provisioning flow must follow

Diagram: Eight-Step Provisioning Sequence: What Must Run Before the First Interaction. Visualizes: Visualize the deterministic eight-step pipeline that runs when a user hits submit on a signup form: Auth → Tenant → Budget → Session → Sandbox → LLM…

When a user hits submit on a signup form, a defined, ordered sequence of eight steps runs before that agent can take its first action: Auth, Tenant, Budget, Session, Context, Sandbox, LLM, and Persist. Auth comes first: the new user's identity gets verified before any resource is allocated at all, because a provisioned VM with no confirmed owner is a liability sitting on the fleet. Tenant creates the tenant record and assigns the user's agent a stable identifier, the key every other component will use to tie resources back to the right person. Budget sets per-user limits, compute ceiling, memory allocation, token budget, before the sandbox ever starts, since enforcing a limit after the fact means killing a running VM and losing its state. Session writes a session record to durable storage so that if the VM crashes mid-provisioning, the orchestrator resumes from the last completed step. Sandbox restores the VM from the golden snapshot; at this point the agent has compute, a filesystem, and a kernel, but none of the user's specific state loaded yet. LLM wires the agent's model endpoint configuration into the VM environment, API keys, routing policy, fallback configuration, so the first turn of conversation doesn't fail on a missing environment variable. Persist writes the completed provisioning state to durable storage and sends a ready signal to the gateway, and only at that point does the agent become reachable by actual user traffic.

Not every step waits on every other step. Auth through Context can run asynchronously and in parallel with each other. Sandbox restore has to wait on Tenant and Budget completing. LLM wiring has to wait on Sandbox. Persist waits on everything. Five independently scalable services carry this pipeline in production: a Gateway for ingress, an Orchestrator, a Scheduler, a Memory service, and a Sandbox service, each one responsible for specific steps and each debuggable on its own. One more property matters for reliability: the pipeline has to be idempotent. If any step fails partway and the orchestrator retries it, the retry should resume from where it left off, not stand up a second VM, a second tenant record, or a second memory store entry for the same user.

Isolation: why each agent needs its own kernel

Agents that install packages, run browser sessions, and execute arbitrary model-written code need an isolation boundary that matches both the threat they pose and the capability they require, and for that workload, a micro-VM is the right fit, not a shared container. A multi-tenant agent fleet has three options, and they trade off differently. Containers share the host kernel. A buggy or malicious agent can exploit a kernel vulnerability that reaches every other tenant on the box, and routine operations like package installs touch a kernel surface that isn't really the agent's own. gVisor intercepts system calls in user space and offers stronger isolation than a plain container, but it blocks direct PCIe passthrough, which rules it out for any agent that needs GPU access, and it adds overhead to every syscall the agent makes. Firecracker gives each workload a kernel it actually owns, running inside a micro-VM with hardware-level isolation, a boot time around 125 milliseconds, and about 5 MB of memory overhead per VM; its limitation is that VFIO device passthrough for real GPU access sits in an unmerged pull request as of late 2026, so it doesn't yet support that path. For agents that can install software and run arbitrary code on a user's behalf, a kernel they own and can modify is the baseline requirement, not a nice-to-have. For a fleet serving many users at once, letting one user's agent affect another's performance or security is not a posture any production system can accept. The micro-VM boundary is what makes the per-user model viable at scale.

Firecracker's density matters for the cost discussion: it only allocates the physical memory pages a VM actually touches, so a VM configured with a generous memory ceiling consumes far less than that ceiling in practice, which lets a single host oversubscribe memory across many agent VMs at once. That density carries a risk: if many VMs spike toward their full allocation at the same moment, the Linux OOM killer starts terminating VMs to protect the host, so a scheduler managing this fleet has to track real memory pressure continuously, treating the configured limits only as a paper ceiling. One more constraint follows from the GPU limitation above: cloud VMs running with nested virtualization block PCIe passthrough outright, so any agent that needs direct GPU access at near-native performance has to run on bare metal.

Making the per-user economics work: idle cost, snapshot-restore scheduling, and the storage-not-compute pricing model

Per-user agents only make financial sense if an idle agent costs about as much as a file sitting on disk, not as much as a server sitting on. Most agents spend most of their time idle: waiting on a tool to respond, waiting on a human approval, waiting on a timer, or just sitting unused between one session and the next. If every provisioned agent keeps its compute running through all of that idle time, the bill scales with the number of signed-up users rather than with how many of them are actually doing something, which breaks the economics at any real scale. The fix mirrors the provisioning pattern already in place. When an agent goes idle, the harness snapshots its full execution state, memory, call stack, pending continuations, to durable storage, then terminates the VM process entirely, dropping its compute cost to zero. When something wakes it back up, the user returns, a tool responds, a timer fires, the VM restores from that snapshot and resumes in under a second, with its full filesystem, installed packages, restored file handles, browser session cookies, and credential store all intact, exactly where it left off. A sleeping agent's disk image sitting on NVMe storage costs a small fraction of what it costs to keep that same VM running continuously, so the cost model shifts shape: instead of compute multiplied by total time, the bill becomes storage multiplied by total time plus compute multiplied by active time only.

Two more levers sit alongside the scheduler. Firecracker's memory hotplugging lets an agent start with a small memory footprint for a simple query and scale up with no downtime when it needs to process a large file, so the fleet pays only for the memory it actually uses at the moment of peak demand, not for every agent's worst-case task provisioned up front. On the model side, you cut cost the most by routing each subtask to the cheapest model that can handle it, ahead of any prompt-level optimization; the spread between budget and top-end models runs roughly 25 to 40 times on both input and output pricing, so task classification inside the orchestrator has to be a core design decision, not something bolted on later. You also need to know, for every action an agent takes, whose authority it was acting under.

Diagram: Active vs. Idle: How Snapshot-Restore Shifts the Agent Cost Model. Visualizes: Show the contrast between two cost models for a per-user agent fleet.

Agent identity and per-user delegation: issuing credentials that scope what each agent can do on behalf of its user

If an agent acts using the user's own login credentials, or one shared service account across the whole fleet, it breaks the audit trail a compliance team needs, and afterward you can't say which user's authority stood behind a given action. Authentication architecture designed for agents in 2026 flags two mistakes teams keep making at this stage. The first is treating the agent like a service account: giving it one long-lived set of broad permissions erases user-level accountability, since every action the agent takes appears to belong to the agent itself with no record of which user it was acting for. The second is treating the agent as a stand-in for the user by handing it the user's own token: that collapses the distinction between agent actions and human actions in the logs, and it lets the token get used for more than the user ever authorized for that specific task.

The pattern that avoids both failures gives the agent a standing identity of its own, separate from any one user, and pairs it with a per-invocation delegation token that carries the user's authority as a distinct layer on top. The agent authenticates as itself. Every action it takes gets attributed to both the agent and the user who delegated it, and the token's scope is bounded to what that user consented to for that invocation, nothing broader. This is the identity half of what makes per-user provisioning viable, and it matters because the agent is no longer a stateless endpoint. It's a standing, persistent actor with its own filesystem, its own memory, and its own credentials, carrying a record of whose authority it was exercising at every step, giving it the same accountability a human colleague would be expected to carry, built directly into the infrastructure.

Sources

  1. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
  2. How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
  3. Infrastructure for the Agentic Web: Gap Analysis and Architecture from the Agentverse Platform
  4. Firecracker

More in Per-User Agent Patterns