DHP – agent work that survives kill -9 DHP (Durable Handoff Protocol), released by developer SIDDARTHAREDDY8 as the pip package dhp-protocol, separates work ownership from worker identity so that agent tasks survive a worker being kill -9'd, OOM-killed, or losing a spot instance mid-run. A supervisor declares a silent worker dead in about 6 seconds, orphans its handoff, bumps a monotonic fence token to reject the dead worker's late writes, and a standby resumes from the last checkpoint with zero rework. DHP positions itself as the durability layer alongside MCP for agent-to-tool calls and A2A for agent-to-agent messaging, and installs via pip or a one-line curl script with MCP registration for Claude Code, Cursor, Windsurf, and Zed. Agent work that survives worker death. When a worker is kill -9 'd, OOM-killed, or its spot instance is reclaimed mid-run, DHP hands the work to a standby that resumes from the last checkpoint — zero rework, zero lost results. MCP standardized agent↔tool. A2A standardized agent↔agent. DHP is the missing layer: agent↔time — execution state that outlives any single worker. pip install dhp-protocol One-line install pip + MCP registration + data dir : curl -fsSL https://raw.githubusercontent.com/SIDDARTHAREDDY8/dhp/main/install.sh | bash Direct into your agent: | Tool | Command | |---|---| | Claude Code | claude mcp add dhp -- dhp-mcp --root ~/.dhp | | Cursor | one-click install https://cursor.com/en/install-mcp?name=dhp&config=eyJjb21tYW5kIjogImRocC1tY3AiLCAiYXJncyI6IFsiLS1yb290IiwgIn4vLmRocCJdfQ | | Windsurf / Zed | copy the JSON from MCP SETUP.md https://github.com/SIDDARTHAREDDY8/dhp/blob/main/MCP SETUP.md | | Any MCP client | dhp-mcp --root ~/.dhp stdio | Other ways to run: Docker https://github.com/SIDDARTHAREDDY8/dhp/blob/main/Dockerfile docker compose up -d for supervisor + transport server · from source git clone + pip install -e ". dev " You delegate a 20-minute task to an AI agent. At minute 14, the process is killed — OOM, spot reclaim, deploy, laptop lid. All 14 minutes of progress are gone. The agent restarts from zero, redoes everything, burns tokens twice. Every agent framework today has the same hole: work is owned by the worker, not the handoff. When the worker dies, the work dies with it. Retries start from scratch. There is no primitive for "this unit of work exists independently of whoever is executing it right now." DHP is that primitive. DHP Durable Handoff Protocol separates work ownership from worker identity : - Work is dispatched as a content-addressed handoff — a durable record of intent task, inputs, output schema, budget . - Workers claim handoffs, hold time-bounded leases , and checkpoint progress as they go. - A supervisor watches leases. A silent worker is declared dead in ~6 seconds; its handoff is orphaned and a monotonic fence token is bumped — the dead worker's late writes are rejected, so two workers can never both believe they own the work. - Any standby recovers the orphaned handoff from the last checkpoint. Zero rework. - Checkpoints ship over TCP as they're made , so if the entire host dies, the peer already has everything. It composes with the existing stack rather than replacing it: | Layer | Protocol | Concern | |---|---|---| | Tools | MCP | agent↔tool calls | | Agents | A2A | agent↔agent messaging | | Durability | DHP | work survives worker death | 1. No lost results — every completed handoff's result is durable and retrievable. 2. No crossed inputs — handoff IDs are content-addressed over intent; inputs cannot be swapped mid-flight. 3. No untyped results — completions are validated against a JSON Schema before acceptance. 4. No silent budget overrun — hard caps on steps, spend, wall-clock, and attempts. Exhausted work parks in a dead-letter queue: intact, visible, never silently dropped. 5. No runner lock-in — suspend → ship → resume on any host, any runner. The wire format is plain JSON-lines. python import dhp store = dhp.Store "./demo-dhp" 1. Dispatch durable work hid = dhp.dispatch store, task kind="demo.count", inputs={"target": 50}, output schema={"type": "object", "properties": {"counted": {"type": "integer"}}, "required": "counted" }, 2. Work it — the context manager heartbeats automatically in the background with dhp.claim store, hid, worker id="w1" as ctx: state = dhp.last checkpoint store, hid or {"n": 0} while state "n" < 50: state "n" += 1 ctx.checkpoint state every step is a resume point ctx.complete {"counted": state "n" } Now kill the worker mid-run and watch recovery: terminal 1: supervisor watches leases, orphans the silent in ~7s dhp-supervisor ./demo-dhp sup1 3600 terminal 2: any standby picks up exactly where the dead worker stopped with dhp.recover store, hid, worker id="w2" as ctx: state = dhp.last checkpoint store, hid {"n": 20} — resumes here while state "n" < 50: state "n" += 1 ctx.checkpoint state ctx.complete {"counted": state "n" } Steps already done are never recomputed . See QUICKSTART.md https://github.com/SIDDARTHAREDDY8/dhp/blob/main/QUICKSTART.md for the full 5-minute walkthrough. A worker holds a handoff under a time-bounded lease default: 2s heartbeat interval, 3 misses ≈ 6s TTL . Heartbeats run in a background thread — slow LLM reasoning between checkpoints never false-orphans a healthy worker. When the supervisor sees an expired lease, it orphans the handoff: ownership moves to a tombstone holder and a monotonic fence token increments. Every subsequent write carries the worker's fence token; a stale worker's late checkpoint is rejected with LeaseLost . This is the same fencing pattern used by Chubby, ZooKeeper, and etcd — exactly-once ownership without distributed consensus. Ten simultaneous recoverers → exactly one wins atomic compare-and-set . The handoff ID is a SHA-256 over the intent — task kind, inputs, output schema, budget — not the mutable state. A recovered handoff is verifiably the same work it was dispatched as. Tampered envelopes and checkpoints are rejected at ingest. Checkpoints are appended to a JSON-lines log — one envelope header, then one line per checkpoint. Two transports: - File dhp.transport : shared directory or volume. - TCP dhp.net : TransportServer on the peer, NetShipper in the worker. Checkpoints arrive as they're made; the server verifies envelope identity and per-checkpoint hashes. The client also buffers locally, so a network partition loses nothing — it replays on reconnect. Rework bound after any failure: ≤ 1 checkpoint . All durability goes through the StoreBackend interface. The default is SQLite/WAL — zero dependencies, crash-safe, single-writer. Postgres advisory locks + FOR UPDATE SKIP LOCKED give identical CAS semantics is a clean extension, not a rewrite. The demo that proves it — not a simulation: python3 demo/run mad demo.py What happens: 1. Host A : worker fetches 60 real Wikipedia pages , checkpointing + shipping each page over TCP to Host B as it's fetched. 2. At page 20: kill -9 on the worker. Real SIGKILL, mid- urlopen . 3. Supervisor orphans the handoff in 6.5s . 4. Host B ingests the 20 shipped checkpoints from its TCP log. 5. Standby on Host B resumes at page 20 — pages 1–20 are never refetched. 6. 60/60 complete. Real CSV dataset on disk. === MAD DEMO: 60 real pages, kill -9 at 20 === worker-A started pid 4414 kill -9 worker-A at 20 pages supervisor orphaned after 6.5s host B ingested: 20 checkpoints, 0 skipped host B resumes from page 20 worker-A died at 20 worker-B started pid 4444 === DONE: 60/60 pages, 60 fetched OK, zero refetched === MAD DEMO: GREEN A network chaos harness demo/net chaos.py SIGKILLs workers mid-ship across rounds: 24/24 checkpoints land on the peer, zero rework. flowchart TB subgraph Interfaces MCP dhp-mcp