Packaging Open-Source Agent Frameworks as Deployable VM Images
VM images give agent frameworks the isolation and reproducibility that library installations cannot.

Open-source agent frameworks can now do far more than the deployment habits around them allow. LangGraph, CrewAI, smolagents, LlamaIndex, and the rest now support orchestrator-worker patterns, evaluator-optimizer loops, and human-in-the-loop workflows that can run for minutes or hours. Yet the dominant way teams put these frameworks into production is still to install them as Python libraries inside a shared runtime, the same way anyone would install a logging utility or an HTTP client. That model has a ceiling, and the ceiling becomes visible the moment the workload stops looking like a simple function call. Dependency conflicts between framework versions turn into coordination problems across teams. An agent that misbehaves can affect every other agent sharing the same kernel, because the shared runtime does not know how to separate one workload's faults from another's. The execution environment itself is never a reproducible artifact; it gets reconstructed, piece by piece, every time something restarts, so you can never quite guarantee that the environment that ran in testing is the one that runs in production. None of this is a flaw in the frameworks. It is a mismatch between what these frameworks now do and the runtime model they have been asked to run inside.
What a VM image boundary gives you that a library installation cannot
When you package a framework as a deployable VM image, what travels with the agent changes. Instead of a set of files installed into someone else's runtime, the isolation boundary, the execution environment, and the agent's state all become part of one artifact, shipped together, restored together. That shift brings three properties with it, and each one answers a specific failure mode from library-based deployment. The first is isolation: each agent gets its own guest kernel, so it can install packages, run shell commands, and execute generated code without touching a shared host kernel or any neighboring agent. The second is persistent state: the working directory, installed dependencies, running processes, and warm memory pages belong to the VM itself, not to some database the application layer has to populate on every invocation. The third is a reproducible environment: the image built at bake time is the exact environment that runs in production, so dependency drift and "it worked locally" failures stop being a category of bug a team has to chase down. The isolation model that makes this practical at scale is the microVM, a lightweight virtual machine built specifically for running many short-lived, untrusted workloads side by side. Firecracker, the best-known implementation, will come up again later in more detail, but the short version is enough for now: it is a small, purpose-built virtual machine monitor, not a general-purpose hypervisor repurposed for this job.
How open-source agent frameworks map onto VM image contents
Every major open-source agent framework has its own runtime shape, meaning its own dependencies, its own execution model, its own surface of state that needs to persist, and that shape decides what belongs inside the VM image and how the image should be built. LangGraph is organized around graph-based control flow, conditional edges, and a checkpointing model, so its VM image needs a persistent checkpoint store baked in: the state machine itself is the artifact worth preserving, and the image is what carries it across invocations. CrewAI takes a role-based approach to multi-agent systems, where each agent is given a defined role, a goal, and a set of tools, and that structure maps naturally onto one VM image per agent, with agents talking to each other across a network boundary. smolagents takes a different approach entirely: its agents write and execute actual Python rather than issuing structured JSON tool calls, so every agent run is a code execution event. That raises the stakes on isolation considerably, since nothing produced by the model can be allowed to share a kernel with any other workload. Frameworks built around retrieval, such as LlamaIndex and Haystack, bring large indexes that belong on the persistent disk layer of the image rather than being re-fetched on every wake-up; once that index is baked in, it survives sleep the same way any other file on disk would. Pydantic AI's typed contracts for inputs, tools, and outputs can be baked into the image as fixed schema fixtures, extending the reproducibility guarantee to the agent's entire input and output contract, not just its code and dependencies. Vendor-native SDKs add a different wrinkle: OpenAI's Agents SDK shipped in March 2025, Google's ADK followed in April 2025, and Anthropic's Agent SDK arrived in September 2025, and each of these pins an agent to one model provider's API surface. When you package that SDK into a VM image, the pin becomes an explicit, auditable dependency that appears in the build, instead of an assumption buried somewhere in a requirements file.
Baking a snapshot: the process that turns a framework install into a deployable image
The mechanism that makes a VM image usable as a unit of deployment is snapshot-baking: run the framework's full startup sequence exactly once, then freeze the resulting machine state as a template other runs can restore from. The sequence itself is straightforward to describe. Install the framework and its dependencies into a fresh microVM guest. Run the agent's actual initialization code, including import resolution, warming up the model client, loading any retrieval index, and running any startup tools the agent depends on. Then snapshot the VM at that point, capturing memory state, disk state, and device state together as one artifact. Restoring from that snapshot later maps the saved memory image back in, loads device state, and unpauses the virtual CPUs, finding the framework's Python interpreter, its imports, its caches, and its heap already sitting there, populated, with the process simply picking up where it was paused. That is a different order of operation from pulling a container image: a container start has to fetch the image, unpack it, mount overlay layers, exec the entrypoint, and run full application initialization, every single time. Snapshot restore skips all of that work because it was already done once, at bake time. The model extends further through forking: a fork restores a specific already-running machine rather than spinning up a fresh copy of the baseline image, and because the forked copies share the parent's memory pages through copy-on-write until they actually diverge, many parallel agent runs branching off one warmed-up template cost far less memory than the same number of agents booted from cold. None of this comes free, though. Networking is the part of a raw Firecracker build that takes real engineering effort: unlike Kubernetes, which has mature CNI plugins to handle pod networking automatically, Firecracker expects the operator to manage TAP interfaces, IP tables, and routing on the host directly. That is genuine operational work, and any team designing a VM image's network layer needs to plan for it.
Persistent state as a first-class artifact, not an application-layer concern
When you move state out of application code and into the VM image boundary, what the agent's memory, working directory, installed tools, and execution context depend on to survive changes. They no longer depend on anyone writing serialization and hydration code; they survive sleep and redeploy because they are simply part of the machine that gets frozen and thawed. The human-in-the-loop pattern makes the gap concrete. On a library-based deployment, pausing a workflow that is waiting on a person means engineering teams have to build dedicated systems to securely persist the state of that workflow, sometimes for long stretches of time, which is application-layer serialization of something a VM already knows how to do on its own. Snapshotting the VM handles this without any of that custom machinery: the working directory, installed dependencies, running processes, and warm page cache are frozen exactly as they stood, and because nothing is running while the VM sits halted, there is no state drifting out of sync between pauses. That same mechanism produces the hibernate-and-wake pattern that makes per-user agents practical: when an agent goes idle, snapshot its VM and halt it; when the user comes back, restore the snapshot, and the agent picks up exactly where it left off, with no cold-start penalty and nothing for application code to reconstruct. This is not a replacement for every application-layer pattern. Sliding-window context management, where an agent writes its reasoning to a scratchpad and compresses older turns into a summary as it approaches a token limit, still belongs at the application layer, and it works alongside VM-level persistence rather than against it: the VM keeps the scratchpad intact between sessions, and the summarization logic decides what actually goes into the model's context window. Files, credentials, browser sessions, and installed packages all live as persistent disk state inside the image, and all of them survive sleep, redeploy, and restart without a line of application code written to bring them back.
Isolation for code-executing agents: why the VM boundary is not optional
Once an agent executes code it produced or retrieved from somewhere else, you are running untrusted code by definition, so the only real question is where the isolation boundary sits. Code written by a model, code pulled from a repository, scripts run by a package manager, browser automation helpers, shell commands the agent generated on its own, and evaluation jobs submitted by end users all show up at execution time carrying behavior nobody has vetted in advance. A microVM gives each agent run its own guest kernel, so if generated code behaves badly, if a package install pulls in something unexpected, or if a command gets executed that should never have touched a shared kernel, the damage stays contained to that one machine. smolagents makes this threat model concrete rather than hypothetical, since its agents write and execute actual Python instead of issuing structured tool calls, so every tool call is a code execution event, and without a VM boundary around it, that execution would share a kernel with every other agent running on the same host. Firecracker's minimalism is what makes the isolation claim credible rather than aspirational: roughly 50,000 lines of Rust implement a small, fixed set of virtio devices, network, block, vsock, balloon, serial console, and a minimal keyboard controller used only to stop the machine, a sharp contrast to the roughly two million lines of C that QEMU uses to emulate a full PC. That smaller surface has been running production workloads at AWS Lambda and Fargate scale since 2018 without a disclosed escape-class vulnerability until 2026, when CVE-2026-5747 and CVE-2026-1386 were disclosed. The isolation guarantee pays off in day-to-day development, too: agents can install packages, break things, and rebuild inside their own VM freely, without any risk of that experimentation touching shared state that other agents or other users depend on.
The economics of idle agents: why the VM image model makes per-user deployment viable
Snapshot-and-restore pushes idle cost toward zero, and that is what turns a dedicated VM image per user from a cost center into a workable product decision. Most agents spend most of their time doing nothing: waiting on a user, waiting on an external event, waiting on the next message in a conversation. If a billing model is built around paying for idle compute, running one VM per user gets expensive at any real scale. But if you pay for storage while that VM sits snapshotted and halted, the same arrangement stays cheap. The mechanism is simple to state: keep a baked snapshot on disk, restore a fresh copy per request, run nothing in between requests, and let idle cost approach zero because nothing is actually consuming CPU time while idle. Every restore starts from the same clean baked baseline, and because copy-on-write memory lets many restores of one template share host RAM, the cost of serving many users off one template stays well below the cost of running that many machines separately. This is the economic case for treating a persistent VM as the natural unit per user: each person gets their own files, credentials, installed tools, browser session, and execution history, without anyone paying for a dedicated machine to sit there running around the clock on their behalf. It helps to keep the billing layers distinct here, because conflating them leads to the wrong conclusions. Model inference, driven by token consumption, remains the dominant cost at scale and has nothing to do with how the agent is packaged. Platform and orchestration costs are a second layer, also largely unaffected by the packaging choice. The VM image approach changes the third layer, sandbox compute, and that is specifically where the idle-cost argument applies. The same structure is what makes multi-tenant agent fleets practical: each tenant's agent runs in its own VM, with its own disk, its own kernel, and its own network identity, so tenant isolation comes from the deployment model itself rather than from code someone has to write and maintain at the application layer.
What the resulting deployment workflow looks like in practice
Once the VM image becomes the actual unit of deployment, the day-to-day workflow looks noticeably different from the old sequence of installing a framework, configuring a runtime, and managing state somewhere off to the side. A developer still starts the same way anyone building an agent would, by picking a framework and writing the agent code locally, so nothing about that first step changes. From there, the work shifts to defining the VM image itself: choosing the guest OS, installing the framework, including whatever tools or packages the agent needs, adding the agent's own code, and deciding what retrieval indexes or credentials belong on the persistent disk layer. Baking the snapshot comes next: boot the VM, run through the full initialization sequence, and snapshot the state that results, producing the one artifact that everything downstream will restore from. That snapshot is then pushed to the infrastructure layer directly, standing in as the deployable artifact itself rather than as an entry in a container registry or a line in a deployment script. What comes back is an API endpoint, and from that point forward, every request restores the snapshotted VM, runs against it, and either returns a result and lets the VM go back to idle, or leaves the VM running so the next message in the conversation picks up exactly where the last one left off. The complexity that used to live scattered across runtime configuration, state management code, and deployment scripts collapses into a single step, building the image, and everything after that step is restore, execute, and either respond or keep waiting.


