{"slug": "show-hn-the-coordination-and-durability-layer-for-multi-agents", "title": "Show HN: The coordination and durability layer for multi-agents", "summary": "CellaFlow launched as a coordination and durability layer for multi-agent AI systems that lets one agent execute an action while the remaining agents subscribe to an in-memory broadcast channel, resolving 50 identical requests with 1 request made. The company reports its embedded RocksDB write-ahead log appends execution steps with 0.31ms marginal engine overhead over bare gRPC, at 1.93ms p50 total commit, and uses composite signatures with Agent_ID and RFC 8785 JSON formatting for persona-isolated caching across Python, TypeScript, and Go. CellaFlow tested four open-source agent products pinned at a commit by killing the process mid-run, finding that a LangGraph agent charging a card before the pod died produced duplicate side effects from the pending node, which a checkpointer does not cover.", "body_md": "The coordination and durability layer for AI agents. When several reach for the same action, one executes, and the rest get the result.\n\nIt charges the card. Then the process dies, the retry starts clean, and it charges again. Or a second agent was already doing it. Or a third reached a different conclusion and acted on it. None of them is wrong, and none of them can see the others.\n\nThat is a coordination problem, and CellaFlow solves it. Agents stay independent; the decision about who acts is taken once, in one place, in front of the action itself.\n\nWhen 50 agents independently scrape the same endpoint, legacy engines fire 50 identical requests, triggering 429 rate limits and multiplying your token bill.\n\nWhen a Researcher and a Coder agent call the same tool with identical inputs, naive caches return the wrong persona's output, corrupting reasoning state.\n\nWhen a stalled agent is timed out and reassigned, the original process keeps running. When it finally wakes up, it writes stale output over fresh results.\n\nWhen Kubernetes reclaims a pod or a spot instance is terminated mid-run, in-memory agent stack context evaporates completely—the swarm halts rather than degrades.\n\nThe first agent executes; the remaining 49 subscribe to an in-memory broadcast channel. 1 request made, all 50 resolved simultaneously.\n\nComposite signatures incorporating the Agent_ID and RFC 8785 JSON formatting ensure persona-isolated and deterministic caching across Python, TS, and Go.\n\nStale writes from superseded zombie processes are deterministically rejected via monotonic storage epochs (`StaleLease` error).\n\nEmbedded RocksDB WAL appends execution steps with just 0.31ms marginal engine overhead over bare gRPC (1.93ms p50 total commit), guaranteeing deterministic replay recovery with zero redundant token spend.\n\nWhen an AI agent retries after an ephemeral compute crash, timeout, or network blip, standard orchestrators replay reasoning from scratch. Without deterministic execution fencing, concurrent workers and retries trigger duplicate payments, duplicate tickets, and irreversible side effects in the real world.\n\nWe took four open-source agent products, pinned them at a commit, and killed the process mid-run. No simulations, no toy loops. This is what their own code does today.\n\nEach one reproduces in about two minutes on a laptop. No API keys and no model spend, because the failure has nothing to do with the model.\n\nA LangGraph agent charges a card, then the pod dies before the checkpoint lands. A second process resumes the thread. The only variable is whether the tool call holds a lease.\n\nThe checkpointer worked correctly in both runs. One seat was reserved each time, because the committed node was skipped on resume. Every duplicate came from the pending node's side effect — the part a checkpointer does not cover.\n\nThe first thing a good engineer reaches for is a row lock or an advisory lock. We built that arm and measured it. It closes every concurrency case and neither crash case, because a lock answers who may act and has nowhere to record what they did.\n\nA Postgres advisory lock frees on connection death, so a worker wedged on a slow model call holds it forever. CellaFlow leases carry a hard lifetime ceiling that is independent of heartbeats, so a zombie holder is reclaimed and the queue behind it drains.\n\nDuplicate work is the easy case. The hard case is two replicas that derived slightly different plans for the same step. CellaFlow sees they are aiming at the same position in the graph and refuses the second before the side effect, rather than serialising two different actions and calling that safe.\n\nThere is one case nothing fixes. If the process dies between the provider acting and anything at all recording it, the work is gone. What a durable record buys is shrinking that window from the whole run down to a single commit.\n\nObserve how Cellaflow journals steps to its local RocksDB store, handles a sudden mid-turn server recycle, and restores execution context in less than 250ms with zero extra LLM tokens.\n\nDurable Tool Call ($0.01 API fee)\n\nLLM Completion Prompt ($0.25 API fee)\n\nDurable Tool Call ($0.05 API fee)\n\nSide-effecting Tool (Idempotent)\n\nNot a queue, not a workflow language, not a rewrite of how your agents think. A single point every agent passes through on its way to the outside world, which knows who is already acting, what has already been done, and what the answer was.\n\nFenced leases with monotonic tokens, so of ten agents reaching for the same action, exactly one proceeds. A hard lifetime ceiling means a holder that hangs without dying cannot block everyone behind it.\n\nThe result and the journal write commit to an append-only execution ledger in a single transaction, so the work survives the process that did it. This is the part a lock cannot copy, and it is why a crash after the side effect is recoverable rather than lost.\n\nA later caller deriving the same key does not run the body. It is handed what the first call returned, across processes, across machines, and across languages. A Python agent and a TypeScript agent converge on one action.\n\nDurability alone gets you the second of these. Coordination is the first and the third, and no amount of retrying gives you those. Because both halves run through one ledger, it can answer a question nothing else can: for a given piece of work, which agent executed and which ones were handed that agent's result.\n\nIn three minutes, a step that charges a customer stops charging them twice.\n\nOne decorator. When your agent's process dies mid-run, the resumed run replays completed steps from the durable log instead of re-executing them — so the charge happens once, not twice.\n\nSingle Docker command — the durable state backend your agents connect to\n\n```\ndocker run -d \\\n  --name cellaflow \\\n  -p 50051:50051 \\\n  -p 9090:9090 \\\n  ghcr.io/cellaflow/cellaflow:latest\n```\n\nVerify it's healthy:\n\n```\ncurl http://localhost:9090/health/ready\n# → {\"status\":\"ready\"}\n```\n\nOn PyPI and npm. No private registry, no tokens needed\n\n```\npip install cellaflow\nnpm install @cellaflow/sdk\n```\n\nCreate `research_agent.py` and run it\n\n``` python\nimport os\nimport sys\nimport uuid\nfrom pathlib import Path\n\nfrom cellaflow import step, tool, workflow\n\nLEDGER = Path(\"charges.log\")\n\n#: Incremented inside the tool body. Stays 0 when the step is replayed, because\n#: a replayed step never enters its body.\nEXECUTED = {\"charges\": 0}\n\n#: Set before the workflow runs so the ledger line can name its session.\nSESSION = {\"id\": \"\"}\n\ndef charges_for(session_id: str) -> list[str]:\n    if not LEDGER.exists():\n        return []\n    return [\n        line\n        for line in LEDGER.read_text().splitlines()\n        if line.startswith(session_id)\n    ]\n\n@tool(tool_name=\"charge_card\")\ndef charge_card(order_id: str, cents: int) -> dict:\n    \"\"\"The irreversible one. Leased, so it happens at most once per session.\"\"\"\n    EXECUTED[\"charges\"] += 1\n    confirmation = f\"ch_{uuid.uuid4().hex[:10]}\"\n    with LEDGER.open(\"a\") as fh:\n        fh.write(f\"{SESSION['id']} {confirmation} {order_id} {cents}\\n\")\n    print(f\"   💳 CHARGED {cents} to {order_id} -> {confirmation}\")\n    return {\"confirmation\": confirmation, \"cents\": cents}\n\n@step\ndef build_receipt(charge: dict, order_id: str) -> dict:\n    print(\"   🧾 Building receipt...\")\n    return {\"order_id\": order_id, \"confirmation\": charge[\"confirmation\"]}\n\n@workflow(version=\"1.0.0\")\ndef checkout(order_id: str, die_after_charging: bool = False) -> dict:\n    charge = charge_card(order_id, 2499)\n\n    if die_after_charging:\n        # The money has moved and nothing durable records the receipt yet.\n        # This is the worst possible moment for the pod to go away.\n        print(\"   💥 pod died\")\n        sys.stdout.flush()\n        os._exit(17)\n\n    return build_receipt(charge, order_id)\n\nif __name__ == \"__main__\":\n    resuming = len(sys.argv) > 1\n    session_id = sys.argv[1] if resuming else str(uuid.uuid4())\n\n    SESSION[\"id\"] = session_id\n\n    if resuming:\n        print(f\"\\n♻️  Resuming session {session_id}\\n\")\n    else:\n        print(f\"\\n▶️  Session {session_id}\")\n        print(\"   (save that id -- you need it to resume)\\n\")\n\n    result = checkout(\n        \"ORD-1001\", die_after_charging=not resuming, _session_id=session_id\n    )\n\n    mine = charges_for(session_id)\n    total = len(LEDGER.read_text().splitlines()) if LEDGER.exists() else 0\n    print(f\"\\n✅ {result}\")\n\n    if resuming and EXECUTED[\"charges\"] == 0:\n        print(\"\\n   charge_card did NOT run -- replayed from the durable log.\")\n        if not mine:\n            print(\n                \"   (No ledger line here for that session: it was charged by a run\"\n                \"\\n    in another directory. The engine still has the record, which\"\n                \"\\n    is why the confirmation above is the original one.)\"\n            )\n    elif resuming:\n        print(\n            \"\\n   ⚠️  charge_card DID run. That session id had no history on this\"\n            \"\\n       engine, so this started a new run rather than resuming one.\"\n            \"\\n       Use the id printed by your own first run.\"\n        )\n\n    print(f\"   this session: {len(mine)} charge(s)\")\n    print(f\"   charges.log:  {total} line(s), every run in this directory\\n\")\n```\n\nSimulate a mid-run pod crash. Resume with the session ID. The customer is charged once.\n\nThat's replay, not suppression. The resumed run receives the result the first run actually committed — it doesn't skip the step and continue with a hole where the confirmation should be.\n\n`_session_id` is create-or-resume, so an unknown id starts a fresh session rather than failing. Use the id your own run printed.\n\nYes — and that's what causes the second charge. If nothing restarted, you'd have one charge and a failed run. The duplicate exists because something retried, and Kubernetes is the retry.\n\nKubernetes restores the process. Your checkpointer restores the state. Neither restores any knowledge of what the process already did to the outside world.\n\ndocs.cellaflow.com/quickstart\n\nCellaFlow replaces heavyweight orchestration clusters with a low-latency, systems-grade engine core designed specifically for cognitive state safety.\n\nCoalesces concurrent tool and LLM calls across thousands of agents. 1 physical execution broadcasts across in-memory Tokio channels to N waiting agents at zero marginal token cost.\n\nNative sub-graph isolation for supervisor-worker patterns. Agents commit atomic JSON deltas with Optimistic Concurrency Control (OCC), avoiding monolithic snapshot bloat.\n\nAn immutable, versioned event ledger that rejects stale zombie writes and guarantees deterministic replay recovery in under 250ms upon container restart.\n\nSingle distroless Docker image with embedded RocksDB storage. 0.31ms marginal engine overhead (~2ms p50 durable commit latency), TLS native encryption, and zero external database dependencies for local and edge deployments.\n\nEvery execution step is journaled with full crash-durability. CellaFlow adds barely ~310µs in marginal engine overhead over raw loopback gRPC transport.\n\nZero-copy serialization of agent state payload into binary format in host memory.\n\nTonic gRPC socket transit, TLS frame packaging, and Bearer token interceptor validation.\n\nCognitive Graph validation, monotonic epoch fencing check, and embedded RocksDB WAL sync.\n\nSingle writer · Embedded RocksDB WAL fsync\n\nWhen an agent commits state, the gRPC socket transport consumes 1.62ms. CellaFlow’s entire stateful persistence engine—including ledger serialization, idempotency validation, and WAL sync—takes only **310 microseconds**.\n\nExternal databases introduce 10–25ms network hops and serialize agent requests under heavy connection pool contention. CellaFlow embeds RocksDB directly inside the 20MB daemon binary.\n\nSee how CellaFlow solves the hardest problems in agentic engineering, from runaway token bills to context collapse.\n\nPrevent 100 autonomous agents from triggering API rate-limit bans or burning duplicate OpenAI tokens during parallel document synthesis.\n\nDrop in `CellaflowSaver` to persist LangGraph state across spot recycles, Lambda timeouts, or container restarts with zero changes to your graph logic.\n\nGuarantee that financial transactions, database writes, and external webhooks execute strictly at most once with deterministic idempotency keys.\n\nRun 30-minute deep codebase refactoring agents safely with Non-Persistable Zones that protect token streams from checkpoint corruption.\n\nSee the immediate business case. Estimate how much LLM API budget you are throwing away on redundant steps due to infrastructure recycles and how Cellaflow's middleware stops the bleed.\n\nEnable background log compaction (5x to 40x compression ratios) to optimize context inputs and save an extra ~45% in prompt token costs.\n\nDirectly thrown away on re-running completed steps.\n\nSaved via asynchronous observation compaction.\n\nAnnual savings from Durable Replay + CCS Memory optimization.\n\nCellaflow maintains a dedicated compilation crate `cellaflow-proto` that compiles Protocol Buffer definitions (`proto/cellaflow/v1/`) on the `cargo build` phase using `tonic-build` in its build script. SDK clients and backend systems share exact runtime contracts without code replication.\n\nTonic gRPC server rejects unencrypted HTTP/2 immediately, securing all pipeline traffic.\n\nBearer Token metadata is validated at the gRPC interceptor layer before passing to the engine.\n\nLocks active executions to the specific version they started on, preventing schema drift.\n\nBare loopback RPC is 1.62ms; CellaFlow's complete atomic WAL journaling and sequence fencing adds only 310µs.\n\n```\nsyntax = \"proto3\";\n\npackage cellaflow.v1;\n\nservice WorkflowEngineService {\n  // Starts a stateful session with registry version pinning\n  rpc StartSession(StartSessionRequest) returns (StartSessionResponse);\n\n  // Commits an execution step with lease validation & sequence guards\n  rpc CommitStep(CommitStepRequest) returns (CommitStepResponse);\n\n  // Recovers full Cognitive Graph history for transparent replay\n  rpc GetGraph(GetGraphRequest) returns (GetGraphResponse);\n\n  // Distributed idempotency cache & non-blocking lease heartbeats\n  rpc CheckIdempotencyCache(CheckCacheRequest) returns (CheckCacheResponse);\n  rpc RenewLease(RenewLeaseRequest) returns (RenewLeaseResponse);\n  rpc ReleaseLease(ReleaseLeaseRequest) returns (ReleaseLeaseResponse);\n}\n```\n\nThe single-node engine is free, forever. One Docker command and a pip install is all you need.\n\nRun the engine locally with Docker, install the Python SDK, and have your first durable workflow running in under 3 minutes.\n\nBuilding something ambitious with multi-agent AI? Book a 30-min technical conversation. We'll review your architecture and help you integrate.\n\nPrefer email? [hello@cellaflow.com](mailto:hello@cellaflow.com)", "url": "https://wpnews.pro/news/show-hn-the-coordination-and-durability-layer-for-multi-agents", "canonical_source": "https://cellaflow.com", "published_at": "2026-10-01 16:06:45+00:00", "updated_at": "2026-10-01 16:18:31.603191+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "developer-tools", "mlops", "ai-safety"], "entities": ["CellaFlow", "LangGraph", "RocksDB", "Kubernetes", "Python", "TypeScript", "Go", "RFC 8785"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-the-coordination-and-durability-layer-for-multi-agents", "markdown": "https://wpnews.pro/news/show-hn-the-coordination-and-durability-layer-for-multi-agents.md", "text": "https://wpnews.pro/news/show-hn-the-coordination-and-durability-layer-for-multi-agents.txt", "jsonld": "https://wpnews.pro/news/show-hn-the-coordination-and-durability-layer-for-multi-agents.jsonld"}}