Networking Multi-Agent Systems Across Self-Hosted Nodes

Distributed agent networks require solving discovery, trust, routing, and state as distinct layers.

Contributing Editor · · 10 min read
Cover illustration for “Networking Multi-Agent Systems Across Self-Hosted Nodes”
Self-Hosting OAF · October 7, 2026 · 10 min read · 2,343 words

A multi-agent system confined to a single host solves coordination through shared memory and a common runtime, but once agents sit on separate self-hosted nodes, that shortcut disappears, and four distinct problems take its place: discovery, trust, routing, and state. Research on distributed general-purpose agent networks frames this shift the way computing history already framed an earlier one. Individual computers did not become valuable by getting faster in isolation; they became valuable by joining an open, stable network that let them cooperate, and agent systems are undergoing the same transition now. What moves across such a network is not a static file or a structured ledger entry. They are semantic declarations, statements about intentions, capabilities, current state, and the constraints under which an agent will cooperate, and that distinction matters because standard peer-to-peer protocols and blockchain-style consensus mechanisms were built for different cargo and do not transfer cleanly onto it. Before any protocol gets chosen, the coordination topology itself, whether hub-and-spoke, pipeline, or peer-to-peer, already fixes much of the system's performance profile and failure modes, making topology a first-class design decision. Each of the four layers that follow has its own failure mode, and most of the breakdowns that occur when a multi-agent prototype moves toward production trace back to solving these layers out of order or treating them as one undifferentiated networking problem.

Discovery: How Agents on Separate Nodes Find Each Other

No trust decision, routing choice, or state handoff can happen until one agent can locate another, which makes discovery the layer every other layer depends on. That makes discovery harder here than in conventional microservice architectures, where service discovery mostly answers "where is it." In an agent network, the question is closer to "what can it do, under what constraints, and is it currently willing." An agent has to announce not just an address but a semantic capability profile: its function, the limits it operates under, and the cooperation rules it will accept from a counterpart it has never interacted with before.

The architecture proposed for distributed general-purpose agent networks treats semantic announcement propagation as one of three core mechanism problems the system has to solve, and the technical route it proposes is bodyless gossip combined with sequential logs. Capability declarations spread through the network the way gossip protocols spread state in distributed systems generally, but they don't carry the full payload each time, so bandwidth use drops while the network stays discoverable. That design choice stands in deliberate contrast to a central registry model. A single registry is a clean idea on paper, but in an open, multi-organization deployment it becomes both a single point of failure and a governance chokepoint, since someone has to control who gets listed and under what rules. Gossip-based propagation gives up some consistency guarantees in exchange for resilience and decentralization, which is the trade a network without one trusted operator generally needs to make.

A usable capability announcement has to carry a handful of things: the agent's identity, the semantics of what it can do, its current state (available, busy, sleeping), the cooperation constraints it operates under, and enough provenance that a receiving agent can decide whether engaging is worthwhile. That last point about state introduces a wrinkle specific to self-hosted deployments running on snapshot-restore infrastructure. Agents that are idle get put to sleep to avoid paying for compute they aren't using, and they wake fast: hundreds of milliseconds for lightweight sandboxes, a few seconds for heavier AI workloads. A sleeping agent is still, in principle, reachable for task dispatch, so the discovery layer cannot simply treat silence as absence. It has to track which nodes are asleep and trigger a wakeup before a message can even be routed to them. Locating an agent, though, only answers whether it exists and what it claims to offer. Whether another agent should actually hand it a task, share context with it, or act on what it says is a separate question entirely, and it is the one trust has to answer.

Establishing trust between agents that don't share a runtime

Trust between agents on different nodes is not a variation on conventional distributed-systems authentication; it is a different problem, because agents can delegate authority to other agents, and that delegation can chain in ways existing security frameworks were not built to reason about. A study of multi-agent system security identifies Identity and Provenance as one of nine distinct risk categories, and names Trust Exploitation as another: attacks that spread through a delegation chain. An attacker does not need to break into the system if an already-trusted agent can be convinced, tricked, or compromised into vouching for something it shouldn't.

That same study found that none of the 16 security frameworks it evaluated achieves majority coverage of any single threat category. Practitioners building distributed agent networks cannot lean on an existing framework to handle inter-agent trust comprehensively; the gap has to be closed architecturally, not by adopting someone else's checklist. The distributed agent network research proposes one technical route for closing part of that gap: BAID-based identity binding paired with MG-EigenTrust reputation scoring. The underlying idea is that trust should not be a binary pass/fail gate. It should accumulate over a history of interactions and stay scoped to topic, so that an agent proven reliable at, say, document summarization is not automatically assumed reliable for something like financial reasoning or code execution.

Compromise of a single agent exposes this problem most sharply around shared memory. When agents on a network read from a common persistent memory store, a single compromised agent can inject false context that degrades the reasoning of every other agent that later reads from that store. That is a genuine attack surface, not a theoretical one, and the MAS security taxonomy's own coverage scores suggest where frameworks currently fall short in watching for it: Memory Poisoning receives the strongest aggregate coverage among the categories studied, at 1.578, while Non-Determinism (1.231) and Data Leakage (1.340) are the most under-addressed. That unevenness means a team can be reasonably well defended against one class of memory attack while remaining exposed on categories that get far less attention from the frameworks it relies on.

This is where isolation choices stop being a performance decision and become a trust decision. Giving each agent its own isolated virtual machine, rather than a shared container or a serverless function sharing a kernel with its neighbors, closes off a specific path: an agent whose runtime gets compromised cannot escape a micro-VM boundary to poison a neighbor's memory or write into a neighbor's filesystem. Trust architecture, understood this way, determines what kinds of multi-agent systems can be built safely at all, not just what gets flagged in a later compliance audit.

Routing messages between agents when network conditions and agent states change

Once an agent has been discovered and judged trustworthy for a given class of task, the network still has to get the right message to it at the right moment, and that moment keeps moving. Routing in a distributed agent network has to track semantics, not just addresses. A request to analyze a document needs to land on an agent whose capability profile actually matches that task, not simply on whichever known address happens to be reachable, since that agent might be occupied, asleep, or poorly suited to the job regardless of its address being valid.

The distributed agent network paper treats open task execution as a mechanism design problem: a leader-follower mechanism-generation loop driven by semantic attribution feedback. The incentive structure of the network has to reward agents for representing their capabilities honestly, because a routing system that can be gamed by agents overstating what they can do will degrade quietly rather than fail loudly, and quiet degradation is far harder to catch in production.

Transport protocol choice has real consequences at this layer. The MCP tooling ecosystem standardized Streamable HTTP as the transport for cloud-hosted agents in the 2025-03-26 spec revision, with further hardening arriving in 2026, while stdio remains the right choice for local and command-line agents, and SSE is treated as legacy for any new deployment. Picking the wrong transport for a given deployment pattern creates a mismatch between how an agent announces itself on the network and how messages actually arrive at it, producing intermittent, hard-to-diagnose delivery failures. The spec revision dated 2026-07-28 goes further, making remote paths stateless: request handling no longer depends on maintaining a protocol session, traffic can be routed by MCP method rather than by session affinity, and clients can cache tool discovery responses using the ttlMs and cacheScope fields that servers are now required to supply. For self-hosted deployments that previously had to maintain sticky sessions across nodes just to keep routing coherent, that statelessness removes a significant source of operational complexity.

The sleeping agent from the discovery layer, which consumes no compute while dormant, still has to be woken and reached here, and that need resurfaces in a sharper form. An agent in snapshot-restore sleep consumes no compute, but it still has to receive messages addressed to it. The routing layer needs an explicit wakeup-and-deliver primitive: restore the agent from its snapshot, then deliver the message, in that order. Wake latency has to be fast enough that the sending agent doesn't time out waiting; sub-second wake for lightweight sandboxes functions as a hard requirement for routing correctness. Topology choice compounds all of this. Hub-and-spoke centralizes routing decisions in one place, which simplifies reasoning about the system but turns that hub into a bottleneck under load. Peer-to-peer spreads routing logic across every node, so you trade bottleneck risk for the requirement that every participant implement routing correctly. Pipeline topology routes simply, but if a single stage fails, it breaks the whole chain behind it.

Keeping agent state coherent across nodes without centralizing it

State is where most production failures in multi-agent networking actually originate, and engineering accounts of running these systems at scale point consistently to state management, not model quality and not the LLM calls themselves, as the layer where things break. If a five-step workflow loses context at step four, it cannot be allowed to restart from step one. It has to resume from the last durable checkpoint, and a system without explicit checkpointing at the orchestration layer has no choice but to rebuild from scratch when that happens, burning compute and very often producing a result that is subtly inconsistent with what came before the failure.

Multi-tenant deployments add a second, non-negotiable constraint on top of checkpointing: one tenant's agent state has to stay invisible to another tenant's agents. That isolation requirement comes from data-compliance obligations, not architectural taste, and it argues for per-agent storage isolation over a shared vector database or a shared memory store that every agent on the fleet can read from. A shared store that is convenient to build is also a shared store that lets the memory poisoning attack described earlier propagate across tenant boundaries, which turns a trust problem into a compliance incident.

The snapshot-restore model that solved part of the discovery and routing puzzle earns its keep again here. When an agent sleeps, its disk and memory state are preserved in that snapshot, so when it wakes, it resumes with full continuity of its prior session. That continuity is the mechanism that makes persistent, coherent state viable across a fleet where most agents spend most of their time idle, because without it, every sleep cycle would force the same kind of costly rebuild that a crashed workflow forces without checkpointing. Isolation and state continuity are the same architectural commitment viewed from two directions: isolating each agent in its own VM is what keeps a compromised neighbor from corrupting shared state, and snapshotting that VM's full disk and memory is what lets state survive the idle periods that make self-hosted economics work.

The Full Four-Layer Stack in a Production Self-Hosted Deployment

A production self-hosted agent fleet is a stack in which discovery, trust, routing, and state each depend on the layer beneath them holding up, and a weakness in one layer surfaces as a mysterious failure in whichever layer sits above it.

The infrastructure requirement that falls out of all four layers together is strong process isolation at the virtual machine level. The trust boundaries and state boundaries this architecture has to enforce, an agent that installs packages, holds credentials, and writes to a persistent disk cannot be contained by a shared-kernel model, because a shared kernel is exactly the thing a compromised agent would use to reach its neighbors. Those trust and state boundaries make VM-level isolation a structural necessity for this kind of network.

The one-OS-user-per-agent pattern gives that requirement a concrete operational form: each long-lived agent gets its own OS user account and its own init-system unit, which maps identity, supervision, and privilege directly onto the host operating system instead of requiring a separate per-agent container orchestration layer on top of it. That pattern was validated in production as of September 2026, and it matters because it keeps the security model legible: an operator can reason about what a given agent can touch by reasoning about what that OS user can touch, using tools the host already provides.

At healthy packing densities on bare metal, self-hosted Firecracker or Kata clusters reach a cost floor that makes deploying one isolated environment per agent economically workable. Where that break-even point falls against rented infrastructure depends on utilization, and the lever that keeps costs in the territory of storage rather than compute is exactly the snapshot-restore idle management described in the discovery and state sections: agents that spend most of their time asleep cost almost nothing while they wait, and wake fast enough to behave, from the network's point of view, as though they were never gone.

What separates a fragile prototype network from a production-grade agent fleet is the decision to treat discovery, trust, routing, and state as one coherent stack, built and reasoned about in that order, rather than as four separate problems bolted together after the fact.

Sources

  1. Security Considerations for Multi-agent Systems
  2. Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and Prototypes
Filed underSelf-Hosting OAF

More in Self-Hosting OAF