cd /news/ai-agents/dhp-agent-work-that-survives-kill-9 · home › topics › ai-agents › article
[ARTICLE · art-148635] src=github.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

DHP – agent work that survives kill -9

DHP (Durable Handoff Protocol), released by developer SIDDARTHAREDDY8 as the pip package dhp-protocol, separates work ownership from worker identity so that agent tasks survive a worker being kill -9'd, OOM-killed, or losing a spot instance mid-run. A supervisor declares a silent worker dead in about 6 seconds, orphans its handoff, bumps a monotonic fence token to reject the dead worker's late writes, and a standby resumes from the last checkpoint with zero rework. DHP positions itself as the durability layer alongside MCP for agent-to-tool calls and A2A for agent-to-agent messaging, and installs via pip or a one-line curl script with MCP registration for Claude Code, Cursor, Windsurf, and Zed.

read9 min views2 publishedOct 10, 2026
DHP – agent work that survives kill -9
Image: Michielbdejong (auto-discovered)

Agent work that survives worker death. When a worker is kill -9'd, OOM-killed, or its spot instance is reclaimed mid-run, DHP hands the work to a standby that resumes from the last checkpoint — zero rework, zero lost results.

MCP standardized agent↔tool. A2A standardized agent↔agent. DHP is the missing layer: agent↔time — execution state that outlives any single worker.

pip install dhp-protocol

One-line install (pip + MCP registration + data dir):

curl -fsSL https://raw.githubusercontent.com/SIDDARTHAREDDY8/dhp/main/install.sh | bash

Direct into your agent:

Tool Command
Claude Code claude mcp add dhp -- dhp-mcp --root ~/.dhp
Cursor one-click install
Windsurf / Zed copy the JSON from MCP_SETUP.md
Any MCP client dhp-mcp --root ~/.dhp (stdio)

Other ways to run: Docker (docker compose up -d for supervisor + transport server) · from source (git clone + pip install -e ".[dev]")

You delegate a 20-minute task to an AI agent. At minute 14, the process is killed — OOM, spot reclaim, deploy, laptop lid. All 14 minutes of progress are gone. The agent restarts from zero, redoes everything, burns tokens twice.

Every agent framework today has the same hole: work is owned by the worker, not the handoff. When the worker dies, the work dies with it. Retries start from scratch. There is no primitive for "this unit of work exists independently of whoever is executing it right now."

DHP is that primitive.

DHP (Durable Handoff Protocol) separates work ownership from worker identity:

  • Work is dispatched as a content-addressed handoff — a durable record of intent (task, inputs, output schema, budget).
  • Workers claim handoffs, holdtime-bounded leases , andcheckpoint progress as they go.
  • A supervisor watches leases. A silent worker is declared dead in ~6 seconds; its handoff isorphaned and a monotonicfence token is bumped — the dead worker's late writes are rejected, so two workers can never both believe they own the work.
  • Any standby recovers the orphaned handoff from the last checkpoint.Zero rework.
  • Checkpoints ship over TCP as they're made , so if the entire host dies, the peer already has everything.

It composes with the existing stack rather than replacing it:

Layer Protocol Concern
Tools MCP agent↔tool calls
Agents A2A agent↔agent messaging
Durability DHP work survives worker death
  1. No lost results — every completed handoff's result is durable and retrievable.
  2. No crossed inputs — handoff IDs are content-addressed over intent; inputs cannot be swapped mid-flight.
  3. No untyped results — completions are validated against a JSON Schema before acceptance.
  4. No silent budget overrun — hard caps on steps, spend, wall-clock, and attempts. Exhausted work parks in a dead-letter queue: intact, visible, never silently dropped.
  5. No runner lock-in — suspend → ship → resume on any host, any runner. The wire format is plain JSON-lines.
import dhp

store = dhp.Store("./demo-dhp")

hid = dhp.dispatch(
    store,
    task_kind="demo.count",
    inputs={"target": 50},
    output_schema={"type": "object",
                   "properties": {"counted": {"type": "integer"}},
                   "required": ["counted"]},
)

with dhp.claim(store, hid, worker_id="w1") as ctx:
    state = dhp.last_checkpoint(store, hid) or {"n": 0}
    while state["n"] < 50:
        state["n"] += 1
        ctx.checkpoint(state)          # every step is a resume point
    ctx.complete({"counted": state["n"]})

Now kill the worker mid-run and watch recovery:

dhp-supervisor ./demo-dhp sup1 3600
with dhp.recover(store, hid, worker_id="w2") as ctx:
    state = dhp.last_checkpoint(store, hid)   # {"n": 20} — resumes here
    while state["n"] < 50:
        state["n"] += 1
        ctx.checkpoint(state)
    ctx.complete({"counted": state["n"]})

Steps already done are never recomputed. See QUICKSTART.md for the full 5-minute walkthrough.

A worker holds a handoff under a time-bounded lease (default: 2s heartbeat interval, 3 misses ≈ 6s TTL). Heartbeats run in a background thread — slow LLM reasoning between checkpoints never false-orphans a healthy worker.

When the supervisor sees an expired lease, it orphans the handoff: ownership moves to a tombstone holder and a monotonic fence token increments. Every subsequent write carries the worker's fence token; a stale worker's late checkpoint is rejected with LeaseLost. This is the same fencing pattern used by Chubby, ZooKeeper, and etcd — exactly-once ownership without distributed consensus.

Ten simultaneous recoverers → exactly one wins (atomic compare-and-set).

The handoff ID is a SHA-256 over the intent — task kind, inputs, output schema, budget — not the mutable state. A recovered handoff is verifiably the same work it was dispatched as. Tampered envelopes and checkpoints are rejected at ingest.

Checkpoints are appended to a JSON-lines log — one envelope header, then one line per checkpoint. Two transports:

  • File (dhp.transport ): shared directory or volume.
  • TCP (dhp.net ):TransportServer on the peer,NetShipper in the worker. Checkpoints arrive as they're made; the server verifies envelope identity and per-checkpoint hashes. The client also buffers locally, so a network partition loses nothing — it replays on reconnect.

Rework bound after any failure: ≤ 1 checkpoint.

All durability goes through the StoreBackend interface. The default is SQLite/WAL — zero dependencies, crash-safe, single-writer. Postgres (advisory locks + FOR UPDATE SKIP LOCKED give identical CAS semantics) is a clean extension, not a rewrite.

The demo that proves it — not a simulation:

python3 demo/run_mad_demo.py

What happens:

  1. Host A : worker fetches60 real Wikipedia pages , checkpointing + shipping each page over TCP toHost B as it's fetched.
  2. At page 20: kill -9 on the worker. Real SIGKILL, mid- urlopen .
  3. Supervisor orphans the handoff in 6.5s .
  4. Host B ingests the 20 shipped checkpoints from its TCP log.
  5. Standby on Host B resumes at page 20 — pages 1–20 are never refetched.
  6. 60/60 complete. Real CSV dataset on disk.
=== MAD DEMO: 60 real pages, kill -9 at 20 ===
worker-A started (pid 4414)
*** kill -9 worker-A at 20 pages ***
supervisor orphaned after 6.5s
host B ingested: 20 checkpoints, 0 skipped
host B resumes from page 20 (worker-A died at 20)
worker-B started (pid 4444)
=== DONE: 60/60 pages, 60 fetched OK, zero refetched ===
MAD DEMO: GREEN

A network chaos harness (demo/net_chaos.py) SIGKILLs workers mid-ship across rounds: 24/24 checkpoints land on the peer, zero rework.

flowchart TB
    subgraph Interfaces
        MCP[dhp-mcp<br/>8 MCP tools]
        SDK[Python SDK<br/>dhp.dispatch/claim/recover]
        A2A[A2A bridge<br/>stable Task ID]
    end
    subgraph Core
        R[Runner<br/>lifecycle: dispatch → claim →<br/>checkpoint → complete]
        S[Supervisor<br/>lease watchdog → orphan<br/>leader election]
    end
    subgraph Durability
        SB[StoreBackend<br/>interface]
        SQ[(SQLite/WAL<br/>default)]
        PG[(Postgres<br/>extension)]
    end
    subgraph Portability
        NET[TCP transport<br/>live checkpoint shipping]
        LOG[JSON-lines log<br/>verified ingest]
    end
    MCP --> R
    SDK --> R
    A2A --> R
    R --> SB
    S --> SB
    SB --> SQ
    SB --> PG
    R --> NET
    NET --> LOG

Module map (dhp/):

Module Responsibility
envelope.py Content-addressed handoff envelopes, identity verification
store.py SQLite/WAL StoreBackend — thread-safe, schema-versioned
backend.py StoreBackend interface (CAS ownership, fences, checkpoints, leadership)
runner3.py Lifecycle: dispatch/claim/recover/checkpoint/complete; TaskContext with auto-heartbeats
supervisor3.py Lease watchdog, orphaning, DLQ, leader election, metrics
transport.py Append-only JSON-lines log: ship/ingest with hash verification
net.py TCP transport: TransportServer +NetShipper , token auth, reconnect
mcp_server.py MCP server (8 tools + health)
a2a_bridge.py A2A durability substrate
config.py /errors.py /validate.py /log.py Tuning, typed errors, input validation, structured logging
import dhp

store = dhp.Store("./data")                        # or any StoreBackend

hid = dhp.dispatch(store, task_kind="...",          # create durable work
                   inputs={...},
                   output_schema={...},
                   budget={"max_attempts": 5})      # optional caps

with dhp.claim(store, hid, worker_id="w1") as ctx: # claim + auto-heartbeat
    state = dhp.last_checkpoint(store, hid) or {}   # resume point
    ctx.checkpoint(state)                           # per chunk
    ctx.heartbeat()                                 # manual (optional w/ ctx mgr)
    ctx.complete(result)                             # schema-validated
    ctx.release("reason")                            # cooperative handoff

with dhp.recover(store, hid, worker_id="w2") as ctx:# recover orphaned work
    ...

dhp.status(store, hid)                              # status/owner/checkpoints
python
from dhp import TransportServer, NetShipper

TransportServer("./peer-logs", port=8471).serve_forever()

shipper = NetShipper("peer.example.com", 8471,
                     local_logpath="./ship-w1.log")  # partition buffer
shipper.ship_envelope(hid, envelope)
shipper.ship_checkpoint(hid, seq, state)             # verified server-side
Command Purpose
dhp-supervisor <root> <id> [seconds] Lease watchdog (orphans the silent)
dhp-mcp --root <dir> MCP server (stdio)
dhp-conform 10 protocol assertions
dhp-chaos Randomized SIGKILL fault injection

dhp_dispatch · dhp_claim · dhp_checkpoint · dhp_heartbeat · dhp_complete · dhp_recover · dhp_status · dhp_last_checkpoint · dhp_health

Not claims — chaos-tested numbers:

Event Measured
kill → orphan → standby resumes ~7s
Double kill (worker + supervisor leader) → recovery 6.6s
Host destroyed → peer ingests TCP log → resumes ~6s
Rework after any kill 0 checkpoints
20-agent MCP soak, 75% worker death rate 20/20 completed
Randomized SIGKILL chaos (14 worker + 5 supervisor kills) all invariants held

Verification: 45 pytest tests (lifecycle, concurrent CAS races, fuzzing, kill -9 integrity, network transport) · dhp-conform 10/10 · network chaos green.

DHP Temporal LangGraph CrewAI / AutoGen
Worker-death detection ✅ leases + supervisor (~7s) ✅ (~12s+) ❌ ❌
Resume from checkpoint, zero rework ✅ ✅ partial (manual) ❌
Cross-host, no shared DB ✅ (TCP log shipping) ❌ (needs cluster) ❌ ❌
Fencing (stale worker rejection) ✅ monotonic tokens ✅ ❌ ❌
Content-addressed work identity ✅ ❌ ❌ ❌
MCP-native ✅ ❌ ❌ ❌
A2A composition ✅ ❌ ❌ ❌
Zero-dependency single node ✅ (SQLite) ❌ (JVM cluster) ✅ ✅

DHP is not a workflow engine — it's the durability primitive workflow engines (and agent frameworks) can build on. If you run Temporal, keep it; DHP is for the agent layer Temporal doesn't reach.

MCP — run dhp-mcp --root <dir> and point any MCP client at it. The skill at skills/dhp-durable-work/SKILL.md teaches agents the dispatch → checkpoint → recover pattern.

A2A — dhp.a2a_bridge makes DHP the durability substrate under A2A tasks: the A2A Task ID stays stable while DHP swaps dead workers underneath (closing A2A's documented "no consensus or global transaction semantics" gap). See BRIDGE_DEMO.md: Agent B killed mid-task → Agent C recovered → client saw WORKING → COMPLETED, the swap invisible.

  • Postgres StoreBackend implementation
  • TLS for the TCP transport
  • OpenTelemetry tracing across handoff attempts
  • dhp dashboard — live handoff/lease/DLQ observability
  • Multi-region supervisor quorum

PRs welcome. The bar: every behavior change needs a chaos test or conformance assertion proving it under kill -9, not just in the happy path.

git clone https://github.com/SIDDARTHAREDDY8/dhp
cd dhp
pip install -e ".[dev]"
python -m pytest tests/ -q   # 45 passed
dhp-conform                   # 10/10
python3 demo/run_mad_demo.py  # the kill demo

MIT — see LICENSE.

── more in #ai-agents 4 stories · sorted by recency
── more on @dhp 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dhp-agent-work-that-…] indexed:0 read:9min 2026-10-10 · —