Self-Hosting AutoGen on Bare Metal vs Managed Cloud

AutoGen entered maintenance mode—teams should build on its successor, the Microsoft Agent Framework.

Staff Writer · · 10 min read
Cover illustration for “Self-Hosting AutoGen on Bare Metal vs Managed Cloud”
Self-Hosting OAF · October 1, 2026 · 10 min read · 2,236 words

Anyone sizing infrastructure for AutoGen in 2026 needs to know a basic fact first: the framework entered maintenance mode in October 2025. As of April 2026, the Microsoft Agent Framework, MAF, is the production-ready successor, and it's the SDK new teams should actually start building on. MAF merges AutoGen with Semantic Kernel, and it arrived at v1.0 with stable APIs, long-term Microsoft support, support for multiple model providers, and cross-runtime interoperability through the A2A and MCP protocols. That merger also raises the floor on what a production deployment requires: Python 3.10 or later, at least 8 GB of RAM in production, and real enterprise configuration work, a heavier baseline than AutoGen itself ever demanded.

None of this makes the question "self-hosting versus managed cloud" obsolete. Teams still running legacy AutoGen or its AG2 community fork face that choice directly. But answering it honestly now requires knowing which side of the fence the reader stands on: planning forward on MAF, or maintaining backward on AutoGen, because the infrastructure consequences of each path diverge in ways that make a single blanket answer misleading.

What each deployment model actually hands back to you (and what it keeps)

Neither self-hosting nor managed cloud eliminates operational burden. Each one distributes it differently, trading specific pieces of control for specific pieces of convenience, and the mistake most teams make is picturing managed cloud as the choice with no burden at all rather than the choice where someone else gets paid to carry it. Self-hosting with frameworks like LangChain, CrewAI, or MAF means building replacements for everything those frameworks assume will already exist in the background: cloud-hosted vector stores, cloud-managed state, cloud-based tracing. The lift involved is bounded rather than boundless.

What self-hosting actually returns to a team is substantial: full control over the operating system, the hardware, and the network; root access to tune databases, install extensions, and wire up custom networking; data that stays inside environments the team controls, with no outside provider sitting in the chain; and kernel-level choices, custom drivers, non-standard operating systems, that no managed platform will ever expose. In exchange, self-hosting keeps every operational chore on the team's own desk: security patches, OS upgrades, hardware failures, performance monitoring, firewall rules, access controls. The entire security posture belongs to whoever runs the hardware.

Managed cloud inverts that trade. It hands back instant provisioning, autoscaling that happens without a phone call, round-the-clock monitoring from the provider, support when something breaks, and no need for in-house staff dedicated to infrastructure. What it keeps from the customer is data sovereignty, since the data lives on the provider's machines; customization below the level of the service's API; predictable billing, since usage-based pricing turns unpredictable under heavy load; and any ability to tune the execution environment at the kernel or driver level. Neither model is the wrong one. They answer different needs at different stages of a company's life, and treating this as a contest with a single winner is how teams make the wrong call for their situation.

Why Agent Workloads Stress Both Models Differently

Agent systems surface infrastructure demands that a typical web application never has to satisfy, and the architectural pattern a team builds toward decides which deployment model can meet those demands. Fan-out parallelization needs compute that scales horizontally without waiting on pre-provisioned capacity. Orchestrator-worker hierarchies need reliable communication between agents and task state that survives across worker lifetimes. Human-in-the-loop patterns, approval gates, pause-and-resume flows, need session lifetimes that last as long as a human takes to decide something, which is fundamentally at odds with sandboxes built to disappear quickly: a sandbox that vanishes mid-deliberation takes the agent's context down with it. Evaluator-optimizer loops need the agent to hold intermediate state across repeated model calls without that state quietly evaporating between them.

Memory is where this pressure concentrates hardest, and it's a place where bare metal has a real advantage over infrastructure billed by the second. Newer approaches to agent memory, such as ontology-driven lifecycle management under the name Fortunate Recall, explicitly model what information should be retained, allowed to decay, merged with other memories, or forgotten, and that kind of modeling needs a durable state store, not a container that disappears at the end of a billing cycle. Self-hosted teams building this kind of memory layer typically reach for Chroma running locally, a self-hosted Qdrant instance, pgvector inside PostgreSQL for hybrid search, or Redis for fast in-memory state and caching, and every one of those options means the team provisions and maintains the store itself.

Security is a major reason teams move toward self-hosting. Agents today routinely carry database credentials, API keys tied to payment processors, and tokens that grant access to cloud infrastructure. A single compromised orchestration platform running multiple tenants puts every one of those customers' secrets at risk at once, while self-hosting keeps the blast radius of any breach confined to a team's own infrastructure.

Firecracker as the shared sandboxing primitive

Firecracker has become the dominant way to sandbox agent code execution no matter which deployment model a team picks, which means the real question isn't whether to use it but who operates it and at which layer of the stack. Built by AWS for Lambda and Fargate, Firecracker is a microVM technology that keeps its attack surface small on purpose: it retains only six virtual devices, virtio-net, virtio-block, a memory balloon, a vsock, a serial port, and a keyboard controller, against the full hardware emulation that something like QEMU provides, and its codebase is a fraction of QEMU's size as a result. Every Firecracker microVM runs its own independent Linux kernel, which closes off the possibility of a kernel vulnerability spreading sideways from one sandbox into another, the isolation property that matters most once a fleet is running multiple tenants side by side. Firecracker boots a microVM in milliseconds, carries minimal memory overhead per instance, and can spin up large numbers of microVMs per second on one host.

Snapshot-restore is what makes this practical at scale: boot a microVM once to a ready state, snapshot its memory and block device to local NVMe storage, and restore future sandboxes from that snapshot instead of booting each one cold. On managed platforms, one widely cited figure puts median sandbox creation time at 78 milliseconds as of January 2026. Independent measurement of self-hosted bare metal running snapshot-restore shows a p50 of 179 milliseconds and a p99 of 203 milliseconds, measured from the moment the guest answers TCP. The gap between those numbers is bounded, and the more consequential variable isn't the raw millisecond difference but who owns and operates the snapshot infrastructure behind it.

AWS Bedrock's AgentCore Runtime is the clearest managed reference implementation of this pattern: it launches lightweight microVMs per session to host MCP servers and agents across multiple tenants, each isolated by session, and passes tenant context, tenant identifier, tier, regional settings, feature flags, entitlements, into those isolated environments through custom HTTP headers. On the self-hosted side, open-source Firecracker-based sandboxing is available with Python and JavaScript SDKs, custom sandbox templates, and an Apache-2.0 licensed runtime a team can run on its own hardware, offered by the same company that also runs a managed version of the same technology. One constraint applies across every managed Firecracker platform regardless of vendor: Firecracker has no PCIe passthrough, so there's no GPU inside the sandbox, which rules out managed Firecracker tiers for any workload step that depends on GPU access.

The three abstraction layers a team can build at, and the time-to-production each one carries

Diagram: The Three Layers of Self-Hosting: Engineering Time vs. Control. Visualizes: Show three distinct abstraction layers a team can build at, arranged vertically by time-to-production (bottom = most effort, top = least).

"Self-hosted" is not one option, and teams that treat it as a single monolithic choice frequently end up building at the wrong layer for what they actually need. There are three distinct layers to choose among, and each carries a different cost in engineering time before anything reaches production.

Building raw Firecracker directly on bare metal puts everything on the team's shoulders: image distribution and caching, networking and routing rules, orchestration, warm pools, autoscaling, observability, and incident response when something fails. Time to production here runs in months, since this is infrastructure engineering in the fullest sense, not application development. In exchange, it buys maximum control over the hardware, direct GPU access, predictable costs once utilization is high, and complete freedom to customize at the kernel level. This layer suits teams that already employ dedicated infrastructure engineers and run workloads large enough to justify the investment.

A self-hosted runtime handles VM lifecycle, snapshot management, and the SDK surface a developer writes against, while the team still owns hosting, networking, and scaling of that runtime. Time to production drops to days, since the hard Firecracker engineering has already been written by someone else. What it buys is transparency into open-source code, portability across hardware, and the ability to run at scale without paying a per-sandbox fee indefinitely.

A fully managed platform sits at the top of the stack. The provider runs everything, and the team's job shrinks to integrating against an SDK. Time to production falls to hours. What it costs in return is billing by the second, caps on session length, a real risk of vendor lock-in, and whatever policy limits the provider places on what an agent is allowed to do inside its sandbox.

For teams building specifically on MAF, self-hosting at any layer still requires Python 3.10 or newer, a minimum of 8 GB of RAM in production, and genuine enterprise configuration work, complexity the framework itself carries regardless of which abstraction layer sits underneath it.

Where the economics flip between renting and owning

Diagram: The Utilization Break-Even: Managed Cloud vs. Bare Metal. Visualizes: Show a single break-even threshold on a utilization axis: below 15–25% utilization, managed cloud wins on pure economics; above 15–25%, bare metal wins.

The break-even point between renting infrastructure and owning it is a question of utilization, not team size or workload category, and most teams never actually run the math to find out where their own number falls. A survey of agent sandbox providers puts that break-even at 15 to 25 percent utilization: below that range, managed cloud wins on pure economics, and above it, bare metal wins.

Managed pricing gives a sense of why utilization matters so much. One major provider charges per vCPU-second and per GiB-second, billed for the entire time a sandbox stays alive rather than only while code is actually executing. Its tier structure runs from a free Hobby tier with a one-time credit and a short maximum session length, through a Pro tier charging a monthly fee for up to 100 concurrent sandboxes (expandable substantially on request) with sessions running up to a full day, up to an Enterprise tier with custom pricing, a substantial monthly minimum, and the option to bring your own cloud account on AWS, GCP, or Azure. A separate provider prices by the core-second at a sandbox-specific rate roughly 3× the standard function rate, with no forced monthly floor on its entry-level plan, though its Team plan does carry a monthly fee, and it is currently the only major sandbox vendor offering GPU access inside the sandbox itself, with a published per-second rate.

The least-discussed variable in this math is idle cost. A real agent loop spends most of its time waiting: the sandbox sits open, the model is reasoning, nothing is executing, and every second of that wait gets billed on usage-based platforms exactly as if work were happening. Bare metal's cost structure runs the opposite way. It's mostly capital expenditure, hardware contracts or server rental fees that cost the same whether agents are active or sitting idle, and that predictability is one of bare metal's strongest advantages over usage-based billing once volume is high enough to justify it. The strongest case against bare metal, though, cuts the other way: cloud wins for workloads that are bursty or run at low duty cycles, and bare metal only wins once utilization is high and stays stable. A team that overestimates its steady-state load ends up owning hardware that sits mostly idle, which flips the entire economic advantage back toward the provider it was trying to avoid.

MCP tooling's hidden cost layer

MCP brings real governance benefits to agent systems, but it also adds a token cost large enough to reshape the economics of a fleet regardless of whether that fleet runs self-hosted or managed. The GitHub MCP server ships with a large number of tools built in, and simply loading those tool definitions into an agent's context window burns tens of thousands of tokens, up to roughly 55,000 depending on configuration and tokenizer, before the agent has answered a single question. That overhead alone can consume roughly half of a typical large-context model's available window.

Benchmarks comparing a command-line interface against MCP tooling found the CLI approach reaching a meaningfully higher task completion score while using roughly the same total number of tokens, and its Token Efficiency Score came out substantially ahead of the MCP approach. That gap matters for infrastructure planning because token overhead doesn't appear as a line item on a cloud bill or a hardware invoice. It accumulates inside the model call itself, and it does so identically whether the agent runs in a self-hosted Firecracker sandbox or a fully managed one. Anyone pricing out a deployment who accounts only for compute, memory, and sandbox creation time is pricing half the system. The remaining cost lives in how much of the context window tool definitions consume before the agent does any actual work, and that cost rides along with every single request the fleet makes, independent of which side of the self-hosting decision a team lands on.

Sources

  1. Self-Hosted vs Managed Cloud: Choosing The Right Infrastructure For Modern Apps - Open Source For You
  2. Self-Hosted AI Agents in 2026: The Complete Guide — Mentiko Blog
  3. Best AutoGen Alternatives (2026): Multi-Agent Framework
  4. AI Agent Code Execution Sandboxes: Isolation from Containers to MicroVMs
  5. How to sandbox AI agents in 2026: Firecracker, gVisor, runtimes & isolation strategies
Filed underSelf-Hosting OAF