{"slug": "dhp-agent-work-that-survives-kill-9", "title": "DHP – agent work that survives kill -9", "summary": "DHP (Durable Handoff Protocol), released by developer SIDDARTHAREDDY8 as the pip package dhp-protocol, separates work ownership from worker identity so that agent tasks survive a worker being kill -9'd, OOM-killed, or losing a spot instance mid-run. A supervisor declares a silent worker dead in about 6 seconds, orphans its handoff, bumps a monotonic fence token to reject the dead worker's late writes, and a standby resumes from the last checkpoint with zero rework. DHP positions itself as the durability layer alongside MCP for agent-to-tool calls and A2A for agent-to-agent messaging, and installs via pip or a one-line curl script with MCP registration for Claude Code, Cursor, Windsurf, and Zed.", "body_md": "**Agent work that survives worker death.** When a worker is `kill -9`'d, OOM-killed, or its spot instance is reclaimed mid-run, DHP hands the work to a standby that resumes from the last checkpoint — zero rework, zero lost results.\n\nMCP standardized agent↔tool. A2A standardized agent↔agent. **DHP is the missing layer: agent↔time** — execution state that outlives any single worker.\n\n```\npip install dhp-protocol\n```\n\n**One-line install** (pip + MCP registration + data dir):\n\n```\ncurl -fsSL https://raw.githubusercontent.com/SIDDARTHAREDDY8/dhp/main/install.sh | bash\n```\n\n**Direct into your agent:**\n\n| Tool | Command | \n|---|---|\n| Claude Code | `claude mcp add dhp -- dhp-mcp --root ~/.dhp` | \n| Cursor | [one-click install](https://cursor.com/en/install-mcp?name=dhp&config=eyJjb21tYW5kIjogImRocC1tY3AiLCAiYXJncyI6IFsiLS1yb290IiwgIn4vLmRocCJdfQ) | \n| Windsurf / Zed | copy the JSON from [MCP_SETUP.md](https://github.com/SIDDARTHAREDDY8/dhp/blob/main/MCP_SETUP.md) | \n| Any MCP client | `dhp-mcp --root ~/.dhp` (stdio) | \n\n**Other ways to run:** [Docker](https://github.com/SIDDARTHAREDDY8/dhp/blob/main/Dockerfile) (`docker compose up -d` for\nsupervisor + transport server) · from source (`git clone` + `pip install -e \".[dev]\"`)\n\nYou delegate a 20-minute task to an AI agent. At minute 14, the process is killed — OOM, spot reclaim, deploy, laptop lid. **All 14 minutes of progress are gone.** The agent restarts from zero, redoes everything, burns tokens twice.\n\nEvery agent framework today has the same hole: **work is owned by the worker, not the handoff.** When the worker dies, the work dies with it. Retries start from scratch. There is no primitive for \"this unit of work exists independently of whoever is executing it right now.\"\n\nDHP is that primitive.\n\nDHP (Durable Handoff Protocol) separates **work ownership** from **worker identity**:\n\n- Work is dispatched as a **content-addressed handoff** — a durable record of intent (task, inputs, output schema, budget).\n- Workers **claim** handoffs, hold**time-bounded leases** , and**checkpoint** progress as they go.\n- A **supervisor** watches leases. A silent worker is declared dead in ~6 seconds; its handoff is**orphaned** and a monotonic**fence token** is bumped — the dead worker's late writes are rejected, so two workers can never both believe they own the work.\n- Any **standby** recovers the orphaned handoff from the last checkpoint.**Zero rework.**\n- Checkpoints ship over **TCP as they're made** , so if the entire host dies, the peer already has everything.\n\nIt composes with the existing stack rather than replacing it:\n\n| Layer | Protocol | Concern | \n|---|---|---|\n| Tools | MCP | agent↔tool calls | \n| Agents | A2A | agent↔agent messaging | \n| **Durability** | **DHP** | **work survives worker death** | \n\n1. **No lost results** — every completed handoff's result is durable and retrievable.\n2. **No crossed inputs** — handoff IDs are content-addressed over intent; inputs cannot be swapped mid-flight.\n3. **No untyped results** — completions are validated against a JSON Schema before acceptance.\n4. **No silent budget overrun** — hard caps on steps, spend, wall-clock, and attempts. Exhausted work parks in a dead-letter queue: intact, visible, never silently dropped.\n5. **No runner lock-in** — suspend → ship → resume on any host, any runner. The wire format is plain JSON-lines.\n\n``` python\nimport dhp\n\nstore = dhp.Store(\"./demo-dhp\")\n\n# 1. Dispatch durable work\nhid = dhp.dispatch(\n    store,\n    task_kind=\"demo.count\",\n    inputs={\"target\": 50},\n    output_schema={\"type\": \"object\",\n                   \"properties\": {\"counted\": {\"type\": \"integer\"}},\n                   \"required\": [\"counted\"]},\n)\n\n# 2. Work it — the context manager heartbeats automatically in the background\nwith dhp.claim(store, hid, worker_id=\"w1\") as ctx:\n    state = dhp.last_checkpoint(store, hid) or {\"n\": 0}\n    while state[\"n\"] < 50:\n        state[\"n\"] += 1\n        ctx.checkpoint(state)          # every step is a resume point\n    ctx.complete({\"counted\": state[\"n\"]})\n```\n\nNow kill the worker mid-run and watch recovery:\n\n```\n# terminal 1: supervisor watches leases, orphans the silent in ~7s\ndhp-supervisor ./demo-dhp sup1 3600\n# terminal 2: any standby picks up exactly where the dead worker stopped\nwith dhp.recover(store, hid, worker_id=\"w2\") as ctx:\n    state = dhp.last_checkpoint(store, hid)   # {\"n\": 20} — resumes here\n    while state[\"n\"] < 50:\n        state[\"n\"] += 1\n        ctx.checkpoint(state)\n    ctx.complete({\"counted\": state[\"n\"]})\n```\n\nSteps already done are **never recomputed**. See [QUICKSTART.md](https://github.com/SIDDARTHAREDDY8/dhp/blob/main/QUICKSTART.md) for the full 5-minute walkthrough.\n\nA worker holds a handoff under a time-bounded lease (default: 2s heartbeat interval, 3 misses ≈ 6s TTL). Heartbeats run in a background thread — slow LLM reasoning between checkpoints never false-orphans a healthy worker.\n\nWhen the supervisor sees an expired lease, it **orphans** the handoff: ownership moves to a tombstone holder and a **monotonic fence token** increments. Every subsequent write carries the worker's fence token; a stale worker's late checkpoint is rejected with `LeaseLost`. This is the same fencing pattern used by Chubby, ZooKeeper, and etcd — exactly-once ownership without distributed consensus.\n\nTen simultaneous recoverers → **exactly one** wins (atomic compare-and-set).\n\nThe handoff ID is a SHA-256 over the *intent* — task kind, inputs, output schema, budget — not the mutable state. A recovered handoff is verifiably the same work it was dispatched as. Tampered envelopes and checkpoints are rejected at ingest.\n\nCheckpoints are appended to a JSON-lines log — one envelope header, then one line per checkpoint. Two transports:\n\n- **File** (`dhp.transport` ): shared directory or volume.\n- **TCP** (`dhp.net` ):`TransportServer` on the peer,`NetShipper` in the worker. Checkpoints arrive as they're made; the server verifies envelope identity and per-checkpoint hashes. The client also buffers locally, so a network partition loses nothing — it replays on reconnect.\n\nRework bound after any failure: **≤ 1 checkpoint**.\n\nAll durability goes through the `StoreBackend` interface. The default is **SQLite/WAL** — zero dependencies, crash-safe, single-writer. Postgres (advisory locks + `FOR UPDATE SKIP LOCKED` give identical CAS semantics) is a clean extension, not a rewrite.\n\nThe demo that proves it — not a simulation:\n\n```\npython3 demo/run_mad_demo.py\n```\n\nWhat happens:\n\n1. **Host A** : worker fetches**60 real Wikipedia pages** , checkpointing + shipping each page over TCP to**Host B** as it's fetched.\n2. At page 20: **`kill -9`** on the worker. Real SIGKILL, mid-` urlopen` .\n3. Supervisor orphans the handoff in **6.5s** .\n4. **Host B** ingests the 20 shipped checkpoints from its TCP log.\n5. Standby on Host B resumes at **page 20** — pages 1–20 are never refetched.\n6. **60/60 complete.** Real CSV dataset on disk.\n\n```\n=== MAD DEMO: 60 real pages, kill -9 at 20 ===\nworker-A started (pid 4414)\n*** kill -9 worker-A at 20 pages ***\nsupervisor orphaned after 6.5s\nhost B ingested: 20 checkpoints, 0 skipped\nhost B resumes from page 20 (worker-A died at 20)\nworker-B started (pid 4444)\n=== DONE: 60/60 pages, 60 fetched OK, zero refetched ===\nMAD DEMO: GREEN\n```\n\nA network chaos harness (`demo/net_chaos.py`) SIGKILLs workers mid-ship across rounds: **24/24 checkpoints land on the peer, zero rework.**\n\n```\nflowchart TB\n    subgraph Interfaces\n        MCP[dhp-mcp<br/>8 MCP tools]\n        SDK[Python SDK<br/>dhp.dispatch/claim/recover]\n        A2A[A2A bridge<br/>stable Task ID]\n    end\n    subgraph Core\n        R[Runner<br/>lifecycle: dispatch → claim →<br/>checkpoint → complete]\n        S[Supervisor<br/>lease watchdog → orphan<br/>leader election]\n    end\n    subgraph Durability\n        SB[StoreBackend<br/>interface]\n        SQ[(SQLite/WAL<br/>default)]\n        PG[(Postgres<br/>extension)]\n    end\n    subgraph Portability\n        NET[TCP transport<br/>live checkpoint shipping]\n        LOG[JSON-lines log<br/>verified ingest]\n    end\n    MCP --> R\n    SDK --> R\n    A2A --> R\n    R --> SB\n    S --> SB\n    SB --> SQ\n    SB --> PG\n    R --> NET\n    NET --> LOG\n```\n\n**Module map** (`dhp/`):\n\n| Module | Responsibility | \n|---|---|\n| `envelope.py` | Content-addressed handoff envelopes, identity verification | \n| `store.py` | SQLite/WAL `StoreBackend` — thread-safe, schema-versioned | \n| `backend.py` | `StoreBackend` interface (CAS ownership, fences, checkpoints, leadership) | \n| `runner3.py` | Lifecycle: dispatch/claim/recover/checkpoint/complete; `TaskContext` with auto-heartbeats | \n| `supervisor3.py` | Lease watchdog, orphaning, DLQ, leader election, metrics | \n| `transport.py` | Append-only JSON-lines log: ship/ingest with hash verification | \n| `net.py` | TCP transport: `TransportServer` +`NetShipper` , token auth, reconnect | \n| `mcp_server.py` | MCP server (8 tools + health) | \n| `a2a_bridge.py` | A2A durability substrate | \n| `config.py` /`errors.py` /`validate.py` /`log.py` | Tuning, typed errors, input validation, structured logging | \n\n``` python\nimport dhp\n\nstore = dhp.Store(\"./data\")                        # or any StoreBackend\n\nhid = dhp.dispatch(store, task_kind=\"...\",          # create durable work\n                   inputs={...},\n                   output_schema={...},\n                   budget={\"max_attempts\": 5})      # optional caps\n\nwith dhp.claim(store, hid, worker_id=\"w1\") as ctx: # claim + auto-heartbeat\n    state = dhp.last_checkpoint(store, hid) or {}   # resume point\n    ctx.checkpoint(state)                           # per chunk\n    ctx.heartbeat()                                 # manual (optional w/ ctx mgr)\n    ctx.complete(result)                             # schema-validated\n    ctx.release(\"reason\")                            # cooperative handoff\n\nwith dhp.recover(store, hid, worker_id=\"w2\") as ctx:# recover orphaned work\n    ...\n\ndhp.status(store, hid)                              # status/owner/checkpoints\npython\nfrom dhp import TransportServer, NetShipper\n\n# peer host:\nTransportServer(\"./peer-logs\", port=8471).serve_forever()\n\n# worker host:\nshipper = NetShipper(\"peer.example.com\", 8471,\n                     local_logpath=\"./ship-w1.log\")  # partition buffer\nshipper.ship_envelope(hid, envelope)\nshipper.ship_checkpoint(hid, seq, state)             # verified server-side\n```\n\n| Command | Purpose | \n|---|---|\n| `dhp-supervisor <root> <id> [seconds]` | Lease watchdog (orphans the silent) | \n| `dhp-mcp --root <dir>` | MCP server (stdio) | \n| `dhp-conform` | 10 protocol assertions | \n| `dhp-chaos` | Randomized SIGKILL fault injection | \n\n`dhp_dispatch` · `dhp_claim` · `dhp_checkpoint` · `dhp_heartbeat` ·\n`dhp_complete` · `dhp_recover` · `dhp_status` · `dhp_last_checkpoint` ·\n`dhp_health`\n\nNot claims — chaos-tested numbers:\n\n| Event | Measured | \n|---|---|\n| kill → orphan → standby resumes | ~7s | \n| Double kill (worker + supervisor leader) → recovery | 6.6s | \n| Host destroyed → peer ingests TCP log → resumes | ~6s | \n| Rework after any kill | **0 checkpoints** | \n| 20-agent MCP soak, 75% worker death rate | 20/20 completed | \n| Randomized SIGKILL chaos (14 worker + 5 supervisor kills) | all invariants held | \n\nVerification: **45 pytest tests** (lifecycle, concurrent CAS races, fuzzing, `kill -9` integrity, network transport) · **`dhp-conform`** 10/10 · **network chaos** green.\n\n|  | DHP | Temporal | LangGraph | CrewAI / AutoGen | \n|---|---|---|---|---|\n| Worker-death detection | ✅ leases + supervisor (~7s) | ✅ (~12s+) | ❌ | ❌ | \n| Resume from checkpoint, zero rework | ✅ | ✅ | partial (manual) | ❌ | \n| Cross-host, no shared DB | ✅ (TCP log shipping) | ❌ (needs cluster) | ❌ | ❌ | \n| Fencing (stale worker rejection) | ✅ monotonic tokens | ✅ | ❌ | ❌ | \n| Content-addressed work identity | ✅ | ❌ | ❌ | ❌ | \n| MCP-native | ✅ | ❌ | ❌ | ❌ | \n| A2A composition | ✅ | ❌ | ❌ | ❌ | \n| Zero-dependency single node | ✅ (SQLite) | ❌ (JVM cluster) | ✅ | ✅ | \n\nDHP is not a workflow engine — it's the durability primitive workflow engines (and agent frameworks) can build on. If you run Temporal, keep it; DHP is for the agent layer Temporal doesn't reach.\n\n**MCP** — run `dhp-mcp --root <dir>` and point any MCP client at it. The skill at [`skills/dhp-durable-work/SKILL.md`](https://github.com/SIDDARTHAREDDY8/dhp) teaches agents the dispatch → checkpoint → recover pattern.\n\n**A2A** — `dhp.a2a_bridge` makes DHP the durability substrate under A2A tasks: the A2A Task ID stays stable while DHP swaps dead workers underneath (closing A2A's documented \"no consensus or global transaction semantics\" gap). See [BRIDGE_DEMO.md](https://github.com/SIDDARTHAREDDY8/dhp/blob/main/BRIDGE_DEMO.md): Agent B killed mid-task → Agent C recovered → client saw WORKING → COMPLETED, the swap invisible.\n\n-  Postgres `StoreBackend` implementation\n- TLS for the TCP transport\n- OpenTelemetry tracing across handoff attempts\n-  `dhp dashboard` — live handoff/lease/DLQ observability\n- Multi-region supervisor quorum\n\nPRs welcome. The bar: every behavior change needs a chaos test or conformance assertion proving it under `kill -9`, not just in the happy path.\n\n```\ngit clone https://github.com/SIDDARTHAREDDY8/dhp\ncd dhp\npip install -e \".[dev]\"\npython -m pytest tests/ -q   # 45 passed\ndhp-conform                   # 10/10\npython3 demo/run_mad_demo.py  # the kill demo\n```\n\nMIT — see [LICENSE](https://github.com/SIDDARTHAREDDY8/dhp/blob/main/LICENSE).", "url": "https://wpnews.pro/news/dhp-agent-work-that-survives-kill-9", "canonical_source": "https://github.com/SIDDARTHAREDDY8/dhp", "published_at": "2026-10-10 06:31:27+00:00", "updated_at": "2026-10-10 07:09:07.157659+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "ai-infrastructure", "developer-tools", "mlops"], "entities": ["DHP", "Durable Handoff Protocol", "SIDDARTHAREDDY8", "MCP", "A2A", "Claude Code", "Cursor", "Windsurf"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/dhp-agent-work-that-survives-kill-9", "markdown": "https://wpnews.pro/news/dhp-agent-work-that-survives-kill-9.md", "text": "https://wpnews.pro/news/dhp-agent-work-that-survives-kill-9.txt", "jsonld": "https://wpnews.pro/news/dhp-agent-work-that-survives-kill-9.jsonld"}}