cd /news/ai-agents/show-hn-how-distributed-claude-code-… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-141454] src=nfltr.xyz β†— pub= topic=ai-agents verified=true sentiment=Β· neutral

Show HN: How distributed Claude Code is built on a tunnelling service

Nfltr is a tunnelling service that distributes Claude Code across multiple machines by turning one Claude session into a hub that spawns and manages agents on joined nodes via tools including spawn_agent, send_message, stop_agent, list_agents, wait_for_agents, list_nodes and start_monitor. Machines join once with `nfltr node join --max-agents 2`, and agents keep running when the hub disconnects, with each completion delivered exactly once so a reattaching hub receives missed results without duplicates. Capabilities are opt-in per machine: agents cannot run shell or `git push` without `--allow-all-tools` (failing with `tools_not_allowed`), monitors require `--allow-monitors`, and nodes clone only repositories listed with `--allow-repo`.

read35 min views2 publishedSep 29, 2026
Show HN: How distributed Claude Code is built on a tunnelling service
Image: source

Claude Code already knows how to split up work. It starts subagents, lets them run in the background, and picks up their results when they finish. That works well as long as everything the work needs is on the machine where Claude Code runs.

Often it is not:

  • The dataset lives on a VM and should not leave it.
  • The failing service is only reachable from inside one network.
  • Three independent fixes would go faster on three machines at once.
  • The machine doing the work gets restarted halfway through.

nfltr keeps Claude Code's model and removes the one-machine limit. Here is how it is built, what went wrong along the way, and how we test it.

The idea: Claude Code's hub, made distributed #

In nfltr, one Claude session is the hub. It gets a small set of tools that mirror Claude Code's own subagent model:

  • spawn_agent starts an agent from a brief and returns its id at once.
  • send_message reaches an agent. A running agent gets it in its live session; a finished one continues in the same session and workspace.
  • stop_agent cancels an agent. Its partial work comes back as evidence.
  • list_agents shows each agent's status, machine, progress and last activity.
  • wait_for_agents blocks until results arrive and returns them.
  • list_nodes shows the joined machines, their descriptions, and how many more agents each may start.
  • start_monitor runs a command on a machine without a model. Each line it prints becomes an event for the hub.

The difference from subagents is where the agents run: on any machine you have joined, and they keep running when the hub goes away.

nfltr orch "<goal>" starts Claude Code as the hub in your terminal. Or add the hub to your own session with claude mcp add nfltr -- nfltr mcp --toolset hub; with nfltr orch hub install-hook, a finished agent wakes an idle session by itself.

Machines join once, with nfltr node join --max-agents 2. A node stays idle until a spawn needs it, then launches a Claude Code agent for that task; the agent exits when the task ends. One join per machine, no configuration per task.

Design principles #

Provide, don't decide

nfltr never picks a machine, a model or a plan for you. A machine or labels named in a spawn are hard filters; among the matches, the runtime only finds a free slot. If several machines could run a monitor, start_monitor asks for one to be named rather than choosing. Model, effort and budgets pass through when set; unset, Claude Code on that machine decides.

The same goes for what we tell the model. In one campaign the hub spawned agents in only 2 of 7 passing runs. Instead of an instruction to delegate, we added facts to the tool results: free capacity, and that agents run in parallel, each with its own workspace and tokens. The choice stays with the hub.

Capabilities are opt-in per machine

What an agent may do on a machine is that machine's decision:

  • Without --allow-all-tools , a node's agents can edit files, but commands that need approval (shell,git push , reading a path outside the checkout) are refused, and the turn fails withtools_not_allowed , naming the flag.
  • Monitors run only with --allow-monitors : they use that machine's credentials.
  • A node clones only repositories it lists with --allow-repo .
  • The API key comes from the environment or the saved config, never from the command line.

Every result arrives exactly once

Agents keep running when the hub disconnects. The hub stores each completion and marks it delivered before wait_for_agents returns it, so a hub that reattaches with the same id gets everything it missed, none of it twice. Monitor events work the same way. In your own Claude Code session the hub id comes from the project directory, so reopening the project reattaches.

Each hub is isolated

A hub sees and controls only its own agents and monitors. A second session in the same project cannot share its hub; it fails and names the holder.

Agents get a clean Claude config

In real-Claude runs, agents picked up the host user's hooks and memory, and one used Claude Code's cross-session messaging to ask another local session for data. Now agents load no user settings, hooks, memory, plugins or MCP servers from the host (the workspace's project context still applies), and cross-session messaging is off. The node verifies both before any task; each is an opt-in per node.

What the relay does, and doesn't #

Every node, agent and hub opens a long-lived connection out to the relay, over TLS, authenticated with your account key. No machine opens an inbound port, joins a VPN or needs a firewall change, so it works behind NAT. The relay routes each message to the agent it is addressed to, stamping the sender from the authenticated connection so nobody can speak as another agent. It runs no model, runs none of your code, and makes no decisions: it never picks a machine, a worker or a retry.

How a result survives a relay restart

Exactly-once delivery is a chain, not one component. The hub keeps its agents' state in its own durable store; each agent keeps its finished result until the hub acknowledges it; every task's events carry a gapless sequence number; and a resumed stream replays what was missed and then says so explicitly.

More of how it works: ownership leases, the capacity feed, idle behaviour, deploys #

  • One owner per task. When processes share a task store (relay replicas, a restarted hub), each task records an owner and an epoch, and each process holds a renewed lease. A task whose owner's lease lapsed can be taken over with the next epoch, and the store rejects any write from a stale owner, so a restart or a second replica never runs a task twice or loses it. A hub restarted afterkill -9 takes the dead process's tasks over at once.
  • No polling for capacity. The relay publishes a small feed of changes (a slot freed, a worker or node connected); a spawn waiting for a machine wakes on the event. With nothing happening, the feed costs nothing.
  • Holding for a reconnecting agent. A dispatch or resume for an agent that is reconnecting is held briefly rather than failed.
  • Artifacts. Where artifacts pass through the relay's store, they are addressed by SHA-256, checked chunk by chunk and as a whole, and resumed from the last acknowledged offset; a mismatch fails closed.
  • Dormant when idle. Periodic work first checks a cheap revision number and reads nothing if it has not changed. Idle after the final hub-on-nodes soak, the relay used 0.23 % CPU (2026-09-28); the production relay idled at about 100 MB and 1 % CPU (2026-09-27).
  • Deploys. nfltr.xyz runs on one VM. A deploy drains, restarts and waits for health before taking traffic; clients ride out the gap for up to 2 minutes, and nodes redial within a second, jittered so a fleet does not reconnect in lockstep. The site is served directly rather than through a CDN proxy, which cut long polls after 100 seconds and answered every request with an error during deploys.

What the relay can see

Connections are encrypted to the relay with TLS. On top of that, the frames between the hub and its agents and nodes (briefs, progress, results, artifacts, monitor lines) are encrypted end to end with X25519 and AES-256-GCM by default, so the relay forwards ciphertext. That is not the whole picture, and we would rather you hear it from us:

  • The dashboard summary is readable by default. The hub also sends the relay a summary of each task for the dashboard: the brief, progress, the result text (up to 12 KiB), steers, usage, branch and commit, and artifact names and hashes. The relay stores it. With--dashboard-digest status the summary carries only state, timing, the machine, usage and counts, and the dashboard says the content stayed on your machine;off sends none. Without one of them, treat the relay as able to read your prompts and results.
  • The end-to-end keys are authenticated only with a pairing key. Without one they protect against a relay that only watches; a relay could in principle sit in the middle of the key exchange, though no code in it does. With--e2ee-key-file , a random secret you copy to your machines yourself, each end proves it holds it before anything is sent, and a relay that swaps keys is refused. It proves the peer is one of your machines, not which one.
  • No downgrade. An end with encryption on refuses unencrypted messages, naming the flag.
  • Metadata is always visible: agent and machine ids, labels and machine descriptions, timing and message sizes.
  • The dashboard can act on your work. The relay carries your own dashboard actions on tasks you started (answer, approve, reject, abort, steer, ) to the hub, unless the hub runs--dashboard-commands=false , which refuses and logs them. It cannot start new work on your machines.

What does not go through the relay: your repositories and data, which stay on the machines that work with them, apart from what agents put in their answers.

How the relay is tested and how reliable it is: relay restarts at random points passed 40 / 40 and 10 / 10 goals (one relay) and 6 / 6 SIGTERM, 3 / 3 SIGKILL and 3 / 3 deploy windows (hub on nodes); ownership leases linearizable over about 630,000 operations; canary 25 / 25 #

The relay is exercised on every rung of the testing ladder: SIGTERM and SIGKILL at random points, deploy windows of 30–90 s, the history audit, the linearizability checks of the ownership leases, and the production canary through the real relay. Relay-side defects those runs found and fixed:

  • The relay kept a context for every finished request on long-lived connections: heap growth of 4.6 MB in 33 minutes, none after the fix.
  • An idle hub re-sent its whole dashboard summary every 5 seconds (58 % of the idle hub's CPU), and the relay's idle CPU went to decoding it. It now sends only on change.
  • Status requests decoded each task's full event history: the relay reached 143.8 % CPU under polling (2026-07-24). A state-only read took 25 polls to 0 full decodes.
  • A relay sharing Redis re-read every task every 5 seconds while idle, for lack of a revision counter.
  • Dispatches and resumes sent while an agent reconnected failed as "target agent not connected"; they are now held briefly. With faster redialing, an in-flight task resumed 0.5 s after a restart instead of 7.5 s (p50).
  • A running agent whose hub had exited, or that had been quiet for over 5 minutes, disappeared from the dashboard; such rows now stay, marked when their publisher is gone.
  • Artifact integrity headers were dropped on the way out of the relay; artifact downloads now bypass the layer that dropped them.

Still open for the relay: multi-replica relays have no soak run yet and no production multi-host proof; the Kubernetes chaos runs used only the scripted hub; the duplicate-execution risk they found is fixed, but its Kubernetes freeze re-run is still to do; rollback has not been run on the real VM.

nfltr vs. Claude Code over SSH or a VPN #

You can already point one Claude Code session at other machines: give it SSH, or put everything on a VPN. That works, and for some jobs it is the right call. The difference is where the agent runs.

One Claude Code session with SSH or a VPN, and nfltr.
Claude Code over SSH or a VPN nfltr
--- --- ---
Where work happens One session drives every machine remotely An agent on each machine, with that machine's tools and project context
Data Command output flows back into one context Processed in place; only results return (the demos return counts, not rows)
Parallel work Background commands, all feeding one session Agents on several machines at once, up to each machine's limit
A dropped connection The running command dies with it The agent keeps going; its result arrives exactly once
Network exposure Inbound SSH or a VPN, keys that reach the whole machine Outbound only; tools, monitors and repositories opt-in per machine
Watching for something The session polls, and every poll costs tokens A monitor costs nothing until a line arrives
Setup and cost None if you already reach the box; one model session A join per machine and the relay; a model session per agent

SSH is the better choice for a quick one-off command on one machine you already reach: nothing to set up, no relay to depend on, and one model session instead of one per agent. A VPN solves reachability, not the agent model: the session still runs in one place and pulls everything back to it.

What we learned building it #

The hub didn't know other machines existed. Given a plain goal about a dataset on a VM, early takes had Claude try gcloud in its own shell. The fix was the instructions the MCP server sends when a session connects: facts about the setup, no policy.

"The QA environment" wasn't covered. The instructions named machines, datasets and services as things that may be elsewhere, but not environments, test suites or logs, and the hub gave up without listing the nodes:

Before: 0 of 3

Goal: "Run the full regression suite now in the QA environment that tests our release branch"

After about 8 seconds, without listing the nodes:

I don't have access to a QA environment

After: 3 of 3

The fact we added to the server's instructions, paraphrased: this machine is one of several; an environment (QA, staging, production), test suite or logs missing here may be on another machine; what is missing here says nothing about the others.

Each run found the release environment and reported the flaky cache test: 34 of 35 passed, and reruns flip-flopped.

"Run it again there once it's back" needs a spawn that can wait. In the resilience demo the VM's node is killed mid-job; the hub sees node_lost and issues the job again. At first a spawn for an absent machine failed after 30 seconds, and the hub announced a "scheduled check" it could not make. Now the spawn's timeout_ms bounds the wait, and it runs when the machine rejoins.

A noisy watcher wakes the session on every line. Each monitor event starts a turn, which is a model call. A command that prints a timestamped "ok" every few seconds wakes it every time, and deduplication cannot match lines that differ; one that prints only changes or failures keeps it quiet. The tool description says so; nfltr does not rewrite lines.

Agent reports were cut to their first line. The result of any agent that changed a repository carried only the first line of its answer:

Before

What the hub received from an agent that pushed a fix branch:

Pushed successfully. Summary:

It asked twice more. That run cost $2.00; the other two runs of the goal cost $0.89 and $1.03.

After

The completions of agents that changed a repository carried their whole answer, 21 and 26 lines in the next runs, and the rerun of that goal needed no follow-up question.

Hubs sharing a store took each other's agents. In a test with no model spend, a second hub on the same machine acknowledged the first hub's parked agents, so their processes exited: every process sharing the one task store acted as a replica of the same hub. Each hub now has its own store and ignores agents not in it.

How we test it #

Most of what goes wrong in a distributed system has nothing to do with the model: a relay restarts mid-dispatch, a machine dies, a laptop sleeps. We test all of that without a model, and spend on real Claude only to see whether it chooses well with what nfltr gives it.

The rules we test by

  • Judge by outcome, not prose. A run passes when the world shows it, not when the model says so.
  • Fix the root cause once, where it lives, with a regression test that fails before the fix.
  • Never widen a timeout to hide a race. A spawn right after a relay restart failed because the node had not re-registered; the fix waits for its registration event within the existing bound.
  • Omission stays omission. A model, effort or budget nobody set is never filled in; a pre-commit check rejects concrete defaults.
  • Tests never run a real agent CLI. A guard puts failing stand-ins forclaude and the other agent CLIs first on each test's path and fails the run if one is called.

A scripted fake Claude

nfltr reaches Claude only through the claude CLI, so a deterministic stand-in for that CLI replaces every model call. It follows a scenario file, answers as an agent or as the hub session, and interprets no prose; anything it does not recognise exits with an error, so drift from the real CLI fails loudly. It proves machinery, not model quality.

End-to-end hub tests with the fake: spawn, steer, stop, continue, reattach, isolation, monitors, upgrade #

These start a real relay, real nodes and the real CLI, and drive the hub over MCP the way Claude Code does: spawns constrained by machine and by repository, a steer into a running agent, stop, a continuation that resumes the same Claude session, a disconnect and a reattach that receives the missed completion exactly once and nothing else, and a kill -9 of the hub whose successor still gets it. Two hubs must not see each other's agents. Monitor lines printed while no hub runs arrive exactly once after reattach, and a runaway yes is held to its declared rate (5 lines delivered, 4.5 million counted as dropped in two seconds). An upgrade test checks that an agent parked in the old shared store is still delivered once. A host-local ladder (readiness, spawn, edit and verify, steer, continue, stop, reattach, on-demand nodes) passes all 8 rungs in 144–172 s on a laptop.

Tests that run a freshly written fake executable used to time out on macOS, which checks a new executable on its first run (0.2–0.9 s idle, seconds under load). The test helper now runs each stub once when it writes it; the product's own bounds stay as they are.

Faults and soaks

With the fake, a soak driver injects one fault per goal at a random point and runs goals back to back for half an hour or more. A goal passes when every agent ends exactly once within 2 minutes of the fault ending (completed, or node_lost where the fault took its node), monitor events arrive exactly once with no gap, neither hub sees the other's agents, the history audit is clean, and nothing is left running. Rates, not gates, from one shared laptop.

The hub on nodes, nine fault classes (2026-09-28)

One relay, three joined nodes launching agents on demand (one runs a monitor printing a tick every second), and two hubs on the same relay. Before is the first runs of the harness; after is the final 30-minute soak.

Fault matrix before and after the fixes, fake Claude, no model spend.
Fault Before After (goals passed)
--- --- ---
Relay SIGTERM or SIGKILL, down 1–5 s fail at spawn the agent exited and the turn hung pass SIGTERM 6 / 6, SIGKILL 3 / 3
Relay deploy window, 30–90 s fail at spawn the turn failed or gave up too early pass 3 / 3
Node kill -9 mid-agent, then rejoin fail the loss arrived after 5 minutes, or never pass 4 / 4, node_lost 5–7 s after the kill
Node restart under a running monitor, hub away fail monitor lines lost; the agent's outcome never arrived pass 2 / 2, every line delivered
Monitor's node killed for good fail the monitor stayed running until the node rejoined pass 3 / 3, lost 30 s after the kill
Hub kill -9 , then reattach fail an agent never exited and kept taking work pass 7 / 7
Node frozen 60–200 s (a closed laptop lid) pass pass 6 / 6
Hub frozen 60–200 s 0 / 6 a stray agent; an empty success in place of a real result pass 2 / 2
Two hubs on one relay pass no cross-hub visibility pass 36 / 36 goals
The six 30-minute soaks, each on the fixes so far. Bar: share of goals passed.
Soak Share passed Goals passed
--- --- ---
Run 1 17 / 24
Run 2 28 / 30
Run 3 28 / 29
Run 4 22 / 25
Run 5 22 / 24
Run 6, final 36 / 36

The final soak ran 36 of 36 goals and 144 agents (134 completed, 10 node_lost on the node a fault took) with no duplicate completion or monitor event and nothing left behind. Idle afterwards, the relay used 0.23 % CPU and the nodes 0.02–0.03 %. The runs found 13 defects; each was fixed at its owner with a regression test.

Recovery in the final soak (p50): relay restart to all nodes back about 1 s, node restart to listed 0.35–0.46 s, hub reattach 114 ms, after a freeze 15 ms (hub) and 356 ms (node) #

Fault Recovery p50 / p95 Measured as
Relay SIGTERM 957 / 1,254 ms relay start to all nodes listed
Relay SIGKILL 1,244 / 1,246 ms same
Deploy window 30–90 s 958 / 2,218 ms same
Node kill -9 and rejoin 354 / 364 ms restart to listed
Node restart under a monitor, hub away 462 / 483 ms same, including the hub's reattach
Hub kill -9 and reattach 114 / 151 ms reattach to first answer
Node frozen 60–200 s 356 / 369 ms resume to listed
Hub frozen 60–200 s 15 / 19 ms resume to answer

Spawn to completion across all goals was 14.1 s p50 and 173.6 s p95; the fake agent's turn is 8 s, and the time includes each fault's own duration.

The 13 defects, in short #

  • A node restart lost the monitor lines it held; they are now kept on disk and delivered after the restart.
  • A node_lost failure could be dropped in transit; it is now kept until the hub acknowledges it.
  • An agent a restarted node had reaped never reported; the hub now notices its agent is gone and ends the turn node_lost .
  • Two cases where a result recorded while the hub was away was never acknowledged, so the agent kept its slot and took new work.
  • A killed node's worktrees stayed behind.
  • Three cases where a status check made up an empty success and dropped the real result (branch, commit, text) the agent still held.
  • A node agent that dropped a not-yet-accepted task during a relay restart exited as idle.
  • A turn refused before it ran ended failed instead of being placed again.
  • A turn waiting for its node after a relay restart gave up just before the node came back: the wait equalled the node's reconnect bound.
  • A monitor on a node that is gone stayed running ; it now endslost , once.

One relay, restarts and long soaks (2026-09-27)

An earlier round, with fleet workers, a node and one hub, restarted the relay at random points and ran a two-hour soak. After its fixes, an in-flight task resumed 0.5 s (p50) after the relay was ready, down from 7.5 s.

Single-relay results and recovery chart: relay restarts 40 / 40 goals (SIGTERM) and 10 / 10 (SIGKILL), node kills 20 / 20, a two-hour soak 303 / 304 #

After a relay restart (SIGTERM runs), seconds from the relay being ready. Top bar: before the follow-up fixes; bottom bar: after.
Measure Before and after Before β†’ after
--- --- ---
In-flight task resumed, p50 7.5 s β†’ 0.5 s
In-flight task resumed, p95 20.4 s β†’ 1.1 s
All workers reconnected, p50 8.7 s β†’ 0.7 s
All workers reconnected, p95 30.7 s β†’ 1.3 s
One relay, fake Claude, no model spend, 2026-09-27.
Fault What is checked Result
--- --- ---
Each hub operation 50 times: spawn (by machine, by repository, with a clone), steer, continue, stop, budget stop, disconnect and reattach passes, no orphan processes or worktrees pass 50 / 50 each
Relay SIGTERM at random points, 57 restarts every agent exactly once, commits on origin, audit clean pass 40 / 40 goals, 120 agents, 0 duplicate or lost
Relay SIGKILL at random points same pass 10 / 10 goals, 17 restarts (first round 9 / 10)
Spawn while the relay is down (deploy window) the spawn waits and runs pass 5 / 5 goals (failed before the fix)
Node kill -9 mid-agent, 20 times the hub gets node_lost ; a new spawn on the rejoined node completes pass 20 / 20, p50 206 ms to node_lost
Hub kill -9 , then reattach missed completion delivered once pass in 30 s (was 111–120 s)
Two-hour soak, 304 goals, 912 agents every goal within its timeout 303 / 304 the miss spanned an 11.6-minute laptop sleep; its agents still completed exactly once
Idle relay and hub for 15 minutes near dormant fixed hub 0 GC/min, ~0.3 % CPU (it had re-read its store every 5 s)

The first restart round found a dispatch abandoned when its stream dropped during a relay restart, a resume refused while the worker reconnected, a node-launched worker that refused its task stranding the agent for about 15 minutes, a dead Claude session's socket mistaken for a live one, and an idle hub that re-read its store every 5 seconds. A second round fixed recovery speed, leftover worktrees, a lost acknowledgement that kept a worker alive, and memory the relay kept for every finished request.

Oracles over what happened

A history audit. In the spirit of a Jepsen checker, a tool reads a run's event history and reports violations, never changing anything: a finished task going live again, a result accepted twice, a gap in a task's event cursor, a completion from the wrong worker, rejected output on the main branch, a running task silent too long.

Linearizability checks. Each task has one owner at a time, fenced by a lease. With porcupine, concurrent replicas run against the in-memory, SQLite and Redis stores under lease expiry, crashes and dropped connections, checked against a model of the contract. After the fixes, 500 seeds (about 630,000 operations) were linearizable for all four store targets (2026-09-25).

What the linearizability checks found: every store flagged at first, four root causes #

A clock read twice in one decision, a clock read before taking the write lock, lease reads not watched to commit, and a retried renewal that extended a lease twice. Each was fixed with a guard test. The 500-seed run included about 1,200 operations with unknown outcomes. A second model over event appends and per-task sequences found no violation.

Kubernetes chaos

A three-node kind cluster runs three relays sharing Redis, workers and a hub. Each experiment is one fault (relay, Redis and worker kills, a rollout restart, partitions, a hub kill; with Chaos Mesh, latency, partitions and clock skew) plus a steady state that must hold before and after, judged from the journal rather than the chaos tool's own status.

The first runs against hub sessions (2026-09-28) used one kind cluster on an arm64 host (10 CPU, 16 GB): three relays on three nodes sharing Redis and the artifact store, two workers, a hub in the cluster, and Chaos Mesh 2.8.4. Each run is a fresh hub session that spawns four agents, each making a commit on its own branch in a 90 s turn; the fault lands while a turn runs. All with the scripted fake Claude, so no model spend. A run passes when every agent completed exactly once, or ended with a typed failure the hub received and could act on, the history audit is clean, and the relays are healthy and idle afterwards. Every experiment passed its 2 valid runs; earlier failing runs are what found the 11 defects below. Clock skew was not validated.

Kubernetes chaos results: 11 experiments, 2 / 2 runs each (22 runs), recovery p50 from 1.0 s to 162 s; clock skew not validated on arm64 #

Three relays on a three-node kind cluster with Chaos Mesh, fake Claude, no model spend, 2026-09-28. Recovery: fault to the first recovery event on the faulted agent's task; for the hub kill, the reattached session's time to finish. "β€”": the running turn never saw the fault and completed on its normal schedule, about 60 s after it.
Fault Recovery p50 Result
--- --- ---
F1 kill the relay that owns the running task 1.4 s; worker and hub reconnect to a surviving relay pass 2 / 2
F2 kill all three relays 8.1 s pass 2 / 2
F3 rollout restart of the relays 12.6 s; the turn is cut when its worker's relay rolls pass 2 / 2
F4 kill Redis β€”; hub tasks live in the hub's own store, not Redis pass 2 / 2
F5 kill a worker mid-commit 1.4 s; the restarted worker takes its task back from its volume pass 2 / 2
F6 freeze a worker past its heartbeat 1.0 s after the thaw; the loss is seen 59–79 s after the freeze, the turn fails at 10 min and the hub retries it pass 2 / 2 (one with a typed failure)
F7 partition the workers from the relays 99 s, about 9 s after the partition heals pass 2 / 2
F8a WAN between workers and relays, 150 Β± 50 ms, 2 % loss β€” pass 2 / 2
F8b latency between relays and Redis β€” pass 2 / 2
F8c partition a relay from Redis β€” pass 2 / 2
F9 kill the hub mid-run, reattach with the same hub id 162 s to finish; the four agents from before the kill ran one turn each and were delivered once pass 2 / 2
Clock skew on a relay and on a worker Chaos Mesh's time fault crashes every Go process it skews on arm64; the harness now refuses it there not validated needs an amd64 cluster

The 11 defects, in short: four in task recovery, two in the fake Claude, five in the harness and deployment #

  • While the backstop for a lost connection waited for a worker to confirm a stop, each resume re-armed it: 21 stops were sent to one worker in 70 ms. It now converges once.
  • A resume that never reached its worker still reopened a failed task, and a thawed worker's first message was recorded before the reopening; the audit saw finished tasks going live again. Such a resume now records nothing, and the reopening comes first.
  • A worker's reconnect report turned a failed task into a completed one without recording that it was taken back; the audit flagged a result accepted twice and a finished task going live again.
  • Every resume armed the task's hard deadline again: 21 timers fired at once and marked a failed task cancelled. There is now one deadline per task.
  • The fake hub ended its session as soon as the frozen turn failed, so the thawed worker's late result reached no one. It now retries a failed agent once.
  • A retried turn whose work the revived first turn had already committed failed on "nothing to commit".
  • The harness looked up the running agent's relay under the wrong key, so some experiments had no relay to fault.
  • Workers could clone the origin before it was seeded, pass their readiness check with an empty repository and refuse every task.
  • Copying the hub's store out failed whenever the store changed during the copy.
  • The pass criterion counted a typed failure the hub can act on as a leak.
  • The clock-skew fault crashed processes on arm64, so a run "passed" as a crash test; it is now refused on arm64 nodes.

Found and fixed: a duplicate-execution risk. Hub attempts carried no link to the attempt that replaced them, so the rule that stops a superseded attempt from being revived could not see that the hub had moved on. In two failing freeze runs, the automatic resume revived failed attempts the hub had already retried and sent one to a worker again under the same task id; it was refused only because the worker was full. Now the hub marks every attempt it moves past (failed, stopped, lost or replaced) as superseded, and each of the ten recovery paths we found refuses to revive, resume or resend a marked attempt; a late result from one is delivered to the hub once, as evidence. A $0 end-to-end test freezes an agent, replaces it on another machine and thaws it, and checks the work ran once. The Kubernetes freeze experiment still has to be re-run against the fix.

Also open: a frozen worker stays listed until its streams drop, so the hub placed two queued agents on it (the typed failure and the retry cover it); only the scripted hub ran, not a real model's recovery choices; two runs per experiment.

A production canary every 15 minutes

A hub on a dedicated VM talks to the hosted relay with real TLS, auth and deploys. Each run starts a fresh hub, spawns a fake agent on each of two canary nodes (one pushes a branch) and a monitor, and requires everything exactly once, the branch on the origin, no other hub's agents in view, and nothing left running. The first 25 runs, over 70 minutes with a relay deploy between two of them, all passed.

The first 25 canary runs (50 agents), laptop hub to the hosted relay, 2026-09-28. Bar: p50; tick: p95; line: max, on a 0–12 s scale. The fake agent's own work is 3 s of spawn to completion.
Step Latency p50 / p95 / max
--- --- ---
spawn_agent call 1.5 / 2.2 / 2.4 s
Spawn to first progress 3.3 / 5.7 / 6.6 s
Spawn to completion 8.5 / 10.9 / 11.6 s
Monitor start to first line 1.1 / 1.6 / 1.9 s

Real Claude, judged by outcome

For the model's choices we run capped campaigns on real machines. Goals are plain English and name no tool; each run is judged by an oracle: statistics recomputed on the VM, branches cloned fresh and tested, a service answering its health check. The latest campaign (2026-09-28): 33 runs on the hosted relay, $14.01 in total, 29 passed. Three failures were one goal before the instruction fix above; the fourth was a wrong count the hub passed on unchecked. No run stalled, exceeded its cap, left anything running, or used a machine the goal ruled out.

The latest real-Claude campaign: 33 runs on the hosted relay, $14.01 in total, each judged by an oracle.
Goal Pass rate Passed
--- --- ---
Statistic on a file that exists only on a VM 4 / 5
Two backlog items in parallel on two machines 3 / 3
Same, laptop only, after the report fix 1 / 1
Broken service on a VM: investigate and fix 3 / 3
Machine killed mid-job, re-run when it is back 3 / 3
Same goal, no fault injected (the test driver missed its trigger) 1 / 1
QA triage, staging: a stopped database 3 / 3
QA triage, main: regression, bisect, fix branch with a test 3 / 3
No machine named: find it from the node descriptions 5 / 5
β€œRun the release suite”, before the instruction fix 0 / 3
Same goal, after the fix 3 / 3

Per-run notes: the wrong count, and two passing runs that still found problems #

The wrong count: the agent compared numbers as strings, and the hub relayed it (model behaviour; the means in the same answer were right). The passing runs: the truncated report (a bug, fixed) and a spawn refused on the laptop for a missing ssh alias, where the error named the cause and the hub moved the work to a VM. We also read every hub transcript and state file afterwards.

We also compare the hub with a single agent on the same tasks, with the same snapshot, goal, oracle and caps. In an earlier comparison on nine scenarios built for tasks where coordinating agents should help, the hub passed 7 and one agent alone passed 5.

What is still open

  • A monitor whose machine is away past the reconnect bound (a laptop asleep for a few minutes) ends lost and is not revived when the machine wakes; the hub can start it again. Lines a killed node never read from the command are gone.
  • After a kill -9 of the hub, the missed completion arrives within 30 s rather than at once.
  • The idle hub is not fully dormant yet (0.36 % CPU, about 5 GC/min in the single-relay soak, possibly partly the sampler).
  • The history audit does not see monitor events; the soak harness checks those itself.
  • A result carrying artifacts that is recorded while the hub was away is acknowledged, but artifact content only the live transfer carries is not delivered on that path.

Demos #

Three of the homepage demos, recorded on real machines:

  • Overnight QA triage. A plain Claude Code session starts a monitor on each of three QA environments and goes idle. When the nightly runs finish, the monitors wake it with nobody typing, and it triages each failure where it happened: a rounding regression pinned to its commit and fixed on a branch with a test, a flaky test rerun, a stopped Postgres started. No logs leave the environments. 677 seconds; session $1.63, seven agents $1.19.
  • An incident across three machines. A monitor wakes the session when a bad deploy starts failing. It investigates on the app VM, runs an aggregate-only query on the private database VM, fixes the code with regression tests, deploys and shows it healthy.

Loaded only when you press play. s over two seconds are shortened; nothing else is edited. Models and budgets on screen were picked for the recording.

Try it #

On each machine:

curl -fsSL https://nfltr.xyz/install.sh | sh
nfltr config add-api-key <key>
nfltr node join --max-agents 2 --allow-all-tools

Then give Claude a goal:

nfltr orch "<goal>"

Or use your own Claude Code session:

claude mcp add nfltr -- nfltr mcp --toolset hub
nfltr orch hub install-hook
claude

More: Use nfltr from Claude Code, Join machines as nodes, and the hub tools reference.

What's next #

Planned, not shipped:

  • An always-on hub. Run the hub on a machine that stays up, so overnight work keeps being handled while your laptop sleeps, and reattach to it from your laptop in the morning.
  • Triggers from external systems. Alertmanager, Jenkins or GitHub webhooks feeding the same inbox as completions and monitor events, with the same limits, their text marked as untrusted data.
  • Schedules for overnight and recurring work. Today a monitor whose command is a loop does this; a scheduler on the node that emits an event on each tick is the next step.
  • A node-scoped key , before nodes run on hosts you do not fully trust, such as PR runners or shared CI.
  • Even less for the relay to see. Per-machine keys (which machine, not only one of yours), and a dashboard view that shows your content without the relay reading it. The status-only dashboard, the pairing key, refusing unencrypted messages and the command switch have shipped; seeWhat the relay can see .
── more in #ai-agents 4 stories Β· sorted by recency
── more on @nfltr 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/show-hn-how-distribu…] indexed:0 read:35min 2026-09-29 Β· β€”