Staff Writer
Published August 24, 2026 · Updated October 7, 2026last updated dates

Google Quietly Dropped SAM: A Peer to Peer Mesh for AI Agents
Google published a small open source project called SAM, short for Sovereign Agent Mesh. It lets AI agents discover each other and call tools across a peer to peer network with no central server in the middle. It reads like infrastructure trivia. It is actually a clue about where agent systems are headed next.
A post by Hasan Toor on X flagged it first: Google had quietly put up a project that lets agents auto discover each other, authenticate every packet, and call tools across the mesh without routing through a central broker. Think BitTorrent, but for agent tool calls. The repo is github.com/google/sam, the docs live at sam-mesh.dev, and the whole thing is Apache 2.0 licensed. It is also explicitly not an officially supported Google product, which matters if you plan to bet on it.
I have been running multi agent fleets in production for two years, so this one caught my attention for practical reasons, not hype. The networking layer for agents has been the missing piece in almost every team I advise. Here is what SAM is, how it works, and whether you should care right now.
What SAM actually is, and what it is not
SAM stands for Sovereign Agent Mesh. Google's own description is blunt: Zero Config, Zero Trust, Agentic Network. In plain language:
- Zero Config: nodes discover each other and build the peer to peer network automatically. You do not wire up service registries by hand.
- Zero Trust: every connection, node, and packet is strictly authenticated. Nothing is trusted because it is on the same network.
- Agentic Network: the network is formed by lightweight local clients called
sam-nodethat provide self healing peer to peer connectivity so agents can plug in, communicate, and invoke tools dynamically. - Portability: cryptographic identities are environment agnostic, so a node can move across cloud, laptop, Raspberry Pi, or Android without reissuing identity.
What SAM is not: it is not Google Segment Anything, it is not a model, it is not an agent framework, and it is not a replacement for your orchestrator. If you use LangGraph, CrewAI, AutoGen, or Google ADK, SAM sits underneath them as a transport layer. Your framework still decides what agents do. SAM decides how they find and talk to each other.
If you want the one sentence mental model: SAM gives each agent a local sidecar that speaks Model Context Protocol on localhost, and that sidecar knows how to find and call the same protocol on any other agent's sidecar anywhere on the mesh.
Why agents need a mesh in the first place
Most agent systems today are still hub and spoke. One orchestrator holds the tool list, keeps the state, and fans work out to workers that report back to the center. It works. It also creates three problems that show up the moment you try to scale or distribute the work.
Tools are stuck to one machine. If agent A on your laptop has a useful tool, say a browser automation or a local file indexer, agent B on a cloud VM cannot call it without you exposing an HTTP endpoint, punching a VPN hole, or deploying both to the same cluster.
Discovery is manual. Adding a new worker means updating a config file, restarting the orchestrator, and hoping the registry stays in sync. Removing one dead worker means the same dance in reverse. There is no self healing.
The broker is a bottleneck and a single point of failure. Every tool call pays a round trip through the center, and if the center goes down, every cross agent call stops.
My fleet hits this weekly. I run implementation workers on schedules, controllers that route work, security lanes with different permissions, and validators that reject bad deploys. Today those pieces talk through a central controller and plain REST. It is reliable but tightly coupled and hard to move between environments. A mesh would let me put a reviewer agent on a cheap edge box at home, a builder agent on a cloud GPU, and have them call each other directly without promoting either into a permanent server.
That is exactly the gap SAM is built for. The project README says agents now run across cloud servers, on prem datacenters, laptops, Raspberry Pis, and Android devices. That spread is real, and the old approach of a shared API gateway does not stretch to it.
How SAM works: three pieces and the identity layer
Under the hood SAM is Go, libp2p, and a small amount of glue. There are three binaries and one concept that ties them together.
| Component | What it does | Where it runs |
|---|---|---|
| sam-control-plane | Registry for node identities, authorization policies, and router coordination | One per mesh, usually in Kubernetes |
| sam-router | libp2p bootstrap and relay, forwards data plane traffic between peers | A handful per mesh for resilience |
| sam-node | Local sidecar that agents talk to on localhost via MCP, handles mesh transport and routing | Everywhere an agent runs |
The important idea is the identity. When you join a mesh, your sam-node does an OIDC login through a browser or device code flow and receives a Biscuit token, a cryptographic capability that it stores locally. That token is reused on every restart, so the node keeps the same peer identity whether it moves from your laptop to a cloud VM. Authentication to the router happens over libp2p, and every packet afterwards is authenticated. That is where the zero trust claim comes from: trust travels with the token, not the network.
Once joined, the node exposes a local MCP server at http://127.0.0.1:8080/mcp by default. Your agent connects to that URL the same way it would connect to any MCP server. SAM adds three things on top of plain MCP:
- discover_remote_services, or
find_remote_toolsin the Python SDK, which queries the distributed hash table for peers exposing a given service name - get_mesh_info, which returns known peers, connected peers, and DHT size
- call_remote_tool, which forwards an MCP tool call to a specific peer by its peer ID
So the loop for an agent is: ask my local sidecar what is on the mesh, pick a peer, call its tool. No central broker inspects the payload, and the sidecar handles discovery, routing, and encryption.
The developer experience: join and run in two commands
Getting on the public testnet at bananas.sam-mesh.dev is deliberately boring, which is a good sign for infrastructure.
curl -sL https://sam-mesh.dev/install.sh | bash
sam-node join https://bananas.sam-mesh.dev
SAM_API_TOKEN=my-secret-token sam-node run --bind-addr 127.0.0.1:8080
Join registers the node and stores the Biscuit. Run starts the sidecar and the local MCP endpoint. The CLI opens your browser for OIDC, or prints a device code if you are headless. Docker users get the same flow with port mappings for 5001 udp and 5002 tcp for libp2p plus 8080 for the API. There is also a --join flag on run that combines the two steps the first time and becomes a no op on later restarts.
The agent side is similarly short. If you use Claude Code, Antigravity, VS Code Copilot, or any MCP aware agent, you add the local URL as a remote MCP server and let the agent's SAM skill handle the rest. For a code driven harness, the Python SDK is one pip install from the repo:
pip install ./sam-mcp-python
from sam_mcp.client import SamClient
async with SamClient(server_url="http://127.0.0.1:8080/mcp") as client:
tools = await client.get_tools()
info = await client.call_tool("get_mesh_info", {})
When I wired a test node locally, the pattern that worked was what the docs recommend: let the agent skill drive the setup, then have the agent report back the one time enrollment URL. The rest is normal MCP tool calling, which is why it feels less like learning a new platform and more like getting a mesh aware extension to a protocol you already use.
What you can actually do with it: the warm pool example
The most convincing part of the repo is not the README but the warm agent pool demo. It solves a real cost problem: some agent tools are expensive to warm up but cheap to reuse, like a code reviewer that holds a model in memory or a sandbox that boots a runtime. You do not want to spawn one per request. You do not want one serial instance either. You want a pool of identical, already running workers and a way to hand them out one job at a time.
SAM builds that pool from nothing but plain MCP services:
- Workers expose a single
review_codetool. Stateless and interchangeable. - A manager exposes
acquire_worker,release_worker, andlist_workers. It pollsfind_remote_tools(code-reviewer)every few seconds via the DHT to learn who exists, and tracks free versus busy with leases. It mints a short lived HMAC token on acquire that the worker verifies offline, so no central check is needed per call. - An orchestrator fans a batch of files across the pool: acquire, call remote tool on that peer, release, all concurrently. Each acquire returns a different free worker, so parallel dispatch never collides.
There is no gossip broadcast, no readiness pub sub. Discovery plus leasing is enough when one manager is the dispatcher. The correctness tricks are small and practical: the manager picks and marks busy synchronously so two concurrent acquires never hand out the same peer, leased workers are never evicted on a transient discovery miss, fencing tokens prevent a late release from freeing the wrong lease, and workers themselves return POOL_BUSY if two calls race. Swap code-reviewer for any warm tool, browser, sandbox, embedder, test runner, and the same manager is reusable.
I recognize this pattern because I run it by hand through a controller and a queue. SAM just makes the pool elastic. Scale a worker deployment up or down mid job and the manager picks it up on the next discovery pass without corrupting in flight leases. That elasticity is the feature central orchestrators charge the most to fake.
Where SAM fits, and where it does not
Three questions tell you quickly whether the mesh helps you.
Do your agents live on more than one kind of host? If every agent runs in one Kubernetes cluster as stateless services, a mesh buys you little today. If you have agents on a developer laptop, a CI runner, a customer VPC, or an edge device, the mesh removes real networking pain.
Do you have warm tools worth pooling? If every call is a short LLM turn with no startup cost, central routing is fine. If you hold models, browsers, or sandboxes open across calls, pooling across hosts pays quickly.
Do you need call level privacy? SAM's zero trust model keeps the broker out of payloads. If your current design routes every tool call through an orchestrator that logs or inspects inputs, the mesh gives you a path where peers talk directly and the control plane only handles enrollment and policy.
There are clear limits right now. The project is tagged early alpha and carries the standard Google disclaimer: not an officially supported Google product, not eligible for the vulnerability rewards program. That is not marketing. It means APIs may move, docs may lag, and you should not put a production customer tenancy on the public bananas testnet without a plan to self host your own control plane and routers in Kubernetes, which the repo does document but adds operational cost. The ecosystem around tool sharing is also young. Discovery finds peers by service name, which works for pools of identical workers but does not yet give you rich capability matching, ranking, or payment.
Compare it to the alternatives you might already know. HashiCorp Consul service mesh gives you discovery and mutual TLS for services, but knows nothing about MCP or agent tool semantics. NATS or Redis give you pub sub, but leave identity, NAT traversal, and peer mobility to you. Tailscale gives you a mesh VPN with strong identity, but leaves tool discovery and remote MCP calling to you. SAM is narrowly scoped to agents and MCP, which is its strength and its boundary.
What this means if you are building an agent product now
Think of SAM as the moment BitTorrent gives you a useful analogy. Before BitTorrent, sharing large files meant one server with one bottleneck. BitTorrent made every downloader also an uploader, replaced the central index with a distributed hash table, and let the network heal itself as peers came and went. SAM does the same for agent tools. Every agent that joins makes the tool graph richer, discovery replaces manual registries, and relays keep connectivity alive across NATs without every tool exposing a public port.
Three practical moves are worth making this month, even if you do not adopt the mesh on day one.
Speak MCP locally. Whatever your framework is, make your agents expose and consume tools over MCP on localhost. That one habit is what lets you drop a SAM sidecar in later without rewriting the agents. If your agents still call tools as ad hoc HTTP handlers with custom schemas, standardize now.
Separate the pool from the workers. If any tool costs more than a second to warm up, build a manager that leases workers and let orchestrators acquire before calling. That pattern works with or without SAM today, and it is exactly what SAM automates over the mesh tomorrow. My article on loop to graph engineering and the piece on running agents at scale both land on the same conclusion: the cheapest performance win is a real validator with authority to reject, and the second cheapest is a leased pool.
Probe SAM on a noncritical path. Put two small agents on two different hosts, give one a tool the other needs, join both to bananas, and measure what you actually get: cold start time, NAT traversal success, call latency versus your current broker, and what observability you lose when the center no longer sees every call. The install is one binary and two commands. That is a lunch break experiment with a production grade answer.
Should you use it today
For side projects, internal tools, and research fleets, yes, try it now. The cost is low and the lesson is durable even if the project renames or stalls: design your agents as MCP servers with portable identity and the mesh becomes a deployment choice, not a rewrite.
For customer facing production, treat SAM as research preview. If you love the shape but need guarantees, self host a private mesh rather than building on the public testnet, keep your existing orchestrator as the fallback path, and do not remove central audit logging until you have a replacement that sees what the mesh deliberately hides. The same tension I wrote about in deterministic versus stochastic fleets applies here: decentralize execution, keep governance centralized and slow to change.
The larger signal is worth naming. A year ago the agent conversation was about prompts, then loops, then graphs. The next layer is the network that lets graphs span hosts without rebuilding a platform each time. Google's answer is early and narrow, but the direction is right. If you are planning an agent system that will live longer than one deployment environment, build as if the mesh is coming, whether you run SAM itself or not.
Sources: Google SAM README and docs at sam-mesh.dev and github.com/google/sam including the Quick Start, Agent Integration guide, and Warm Agent Pool use case, plus reporting on the project by Epoch. All accessed August 2026. The original X post that surfaced the project is Hasan Toor's August 22, 2026 thread.
About the Author
I'm Brian Marvin, an AI-native Fractional CTO with 30 years in technical leadership. At empowered.guru, I help startups build MVPs, shape roadmaps, and make AI-powered technology decisions that scale.
