# Show HN: The coordination and durability layer for multi-agents

> Source: <https://cellaflow.com>
> Published: 2026-10-01 16:06:45+00:00

The coordination and durability layer for AI agents. When several reach for the same action, one executes, and the rest get the result.

It charges the card. Then the process dies, the retry starts clean, and it charges again. Or a second agent was already doing it. Or a third reached a different conclusion and acted on it. None of them is wrong, and none of them can see the others.

That is a coordination problem, and CellaFlow solves it. Agents stay independent; the decision about who acts is taken once, in one place, in front of the action itself.

When 50 agents independently scrape the same endpoint, legacy engines fire 50 identical requests, triggering 429 rate limits and multiplying your token bill.

When a Researcher and a Coder agent call the same tool with identical inputs, naive caches return the wrong persona's output, corrupting reasoning state.

When a stalled agent is timed out and reassigned, the original process keeps running. When it finally wakes up, it writes stale output over fresh results.

When Kubernetes reclaims a pod or a spot instance is terminated mid-run, in-memory agent stack context evaporates completely—the swarm halts rather than degrades.

The first agent executes; the remaining 49 subscribe to an in-memory broadcast channel. 1 request made, all 50 resolved simultaneously.

Composite signatures incorporating the Agent_ID and RFC 8785 JSON formatting ensure persona-isolated and deterministic caching across Python, TS, and Go.

Stale writes from superseded zombie processes are deterministically rejected via monotonic storage epochs (`StaleLease` error).

Embedded RocksDB WAL appends execution steps with just 0.31ms marginal engine overhead over bare gRPC (1.93ms p50 total commit), guaranteeing deterministic replay recovery with zero redundant token spend.

When an AI agent retries after an ephemeral compute crash, timeout, or network blip, standard orchestrators replay reasoning from scratch. Without deterministic execution fencing, concurrent workers and retries trigger duplicate payments, duplicate tickets, and irreversible side effects in the real world.

We took four open-source agent products, pinned them at a commit, and killed the process mid-run. No simulations, no toy loops. This is what their own code does today.

Each one reproduces in about two minutes on a laptop. No API keys and no model spend, because the failure has nothing to do with the model.

A LangGraph agent charges a card, then the pod dies before the checkpoint lands. A second process resumes the thread. The only variable is whether the tool call holds a lease.

The checkpointer worked correctly in both runs. One seat was reserved each time, because the committed node was skipped on resume. Every duplicate came from the pending node's side effect — the part a checkpointer does not cover.

The first thing a good engineer reaches for is a row lock or an advisory lock. We built that arm and measured it. It closes every concurrency case and neither crash case, because a lock answers who may act and has nowhere to record what they did.

A Postgres advisory lock frees on connection death, so a worker wedged on a slow model call holds it forever. CellaFlow leases carry a hard lifetime ceiling that is independent of heartbeats, so a zombie holder is reclaimed and the queue behind it drains.

Duplicate work is the easy case. The hard case is two replicas that derived slightly different plans for the same step. CellaFlow sees they are aiming at the same position in the graph and refuses the second before the side effect, rather than serialising two different actions and calling that safe.

There is one case nothing fixes. If the process dies between the provider acting and anything at all recording it, the work is gone. What a durable record buys is shrinking that window from the whole run down to a single commit.

Observe how Cellaflow journals steps to its local RocksDB store, handles a sudden mid-turn server recycle, and restores execution context in less than 250ms with zero extra LLM tokens.

Durable Tool Call ($0.01 API fee)

LLM Completion Prompt ($0.25 API fee)

Durable Tool Call ($0.05 API fee)

Side-effecting Tool (Idempotent)

Not a queue, not a workflow language, not a rewrite of how your agents think. A single point every agent passes through on its way to the outside world, which knows who is already acting, what has already been done, and what the answer was.

Fenced leases with monotonic tokens, so of ten agents reaching for the same action, exactly one proceeds. A hard lifetime ceiling means a holder that hangs without dying cannot block everyone behind it.

The result and the journal write commit to an append-only execution ledger in a single transaction, so the work survives the process that did it. This is the part a lock cannot copy, and it is why a crash after the side effect is recoverable rather than lost.

A later caller deriving the same key does not run the body. It is handed what the first call returned, across processes, across machines, and across languages. A Python agent and a TypeScript agent converge on one action.

Durability alone gets you the second of these. Coordination is the first and the third, and no amount of retrying gives you those. Because both halves run through one ledger, it can answer a question nothing else can: for a given piece of work, which agent executed and which ones were handed that agent's result.

In three minutes, a step that charges a customer stops charging them twice.

One decorator. When your agent's process dies mid-run, the resumed run replays completed steps from the durable log instead of re-executing them — so the charge happens once, not twice.

Single Docker command — the durable state backend your agents connect to

```
docker run -d \
  --name cellaflow \
  -p 50051:50051 \
  -p 9090:9090 \
  ghcr.io/cellaflow/cellaflow:latest
```

Verify it's healthy:

```
curl http://localhost:9090/health/ready
# → {"status":"ready"}
```

On PyPI and npm. No private registry, no tokens needed

```
pip install cellaflow
npm install @cellaflow/sdk
```

Create `research_agent.py` and run it

``` python
import os
import sys
import uuid
from pathlib import Path

from cellaflow import step, tool, workflow

LEDGER = Path("charges.log")

#: Incremented inside the tool body. Stays 0 when the step is replayed, because
#: a replayed step never enters its body.
EXECUTED = {"charges": 0}

#: Set before the workflow runs so the ledger line can name its session.
SESSION = {"id": ""}

def charges_for(session_id: str) -> list[str]:
    if not LEDGER.exists():
        return []
    return [
        line
        for line in LEDGER.read_text().splitlines()
        if line.startswith(session_id)
    ]

@tool(tool_name="charge_card")
def charge_card(order_id: str, cents: int) -> dict:
    """The irreversible one. Leased, so it happens at most once per session."""
    EXECUTED["charges"] += 1
    confirmation = f"ch_{uuid.uuid4().hex[:10]}"
    with LEDGER.open("a") as fh:
        fh.write(f"{SESSION['id']} {confirmation} {order_id} {cents}\n")
    print(f"   💳 CHARGED {cents} to {order_id} -> {confirmation}")
    return {"confirmation": confirmation, "cents": cents}

@step
def build_receipt(charge: dict, order_id: str) -> dict:
    print("   🧾 Building receipt...")
    return {"order_id": order_id, "confirmation": charge["confirmation"]}

@workflow(version="1.0.0")
def checkout(order_id: str, die_after_charging: bool = False) -> dict:
    charge = charge_card(order_id, 2499)

    if die_after_charging:
        # The money has moved and nothing durable records the receipt yet.
        # This is the worst possible moment for the pod to go away.
        print("   💥 pod died")
        sys.stdout.flush()
        os._exit(17)

    return build_receipt(charge, order_id)

if __name__ == "__main__":
    resuming = len(sys.argv) > 1
    session_id = sys.argv[1] if resuming else str(uuid.uuid4())

    SESSION["id"] = session_id

    if resuming:
        print(f"\n♻️  Resuming session {session_id}\n")
    else:
        print(f"\n▶️  Session {session_id}")
        print("   (save that id -- you need it to resume)\n")

    result = checkout(
        "ORD-1001", die_after_charging=not resuming, _session_id=session_id
    )

    mine = charges_for(session_id)
    total = len(LEDGER.read_text().splitlines()) if LEDGER.exists() else 0
    print(f"\n✅ {result}")

    if resuming and EXECUTED["charges"] == 0:
        print("\n   charge_card did NOT run -- replayed from the durable log.")
        if not mine:
            print(
                "   (No ledger line here for that session: it was charged by a run"
                "\n    in another directory. The engine still has the record, which"
                "\n    is why the confirmation above is the original one.)"
            )
    elif resuming:
        print(
            "\n   ⚠️  charge_card DID run. That session id had no history on this"
            "\n       engine, so this started a new run rather than resuming one."
            "\n       Use the id printed by your own first run."
        )

    print(f"   this session: {len(mine)} charge(s)")
    print(f"   charges.log:  {total} line(s), every run in this directory\n")
```

Simulate a mid-run pod crash. Resume with the session ID. The customer is charged once.

That's replay, not suppression. The resumed run receives the result the first run actually committed — it doesn't skip the step and continue with a hole where the confirmation should be.

`_session_id` is create-or-resume, so an unknown id starts a fresh session rather than failing. Use the id your own run printed.

Yes — and that's what causes the second charge. If nothing restarted, you'd have one charge and a failed run. The duplicate exists because something retried, and Kubernetes is the retry.

Kubernetes restores the process. Your checkpointer restores the state. Neither restores any knowledge of what the process already did to the outside world.

docs.cellaflow.com/quickstart

CellaFlow replaces heavyweight orchestration clusters with a low-latency, systems-grade engine core designed specifically for cognitive state safety.

Coalesces concurrent tool and LLM calls across thousands of agents. 1 physical execution broadcasts across in-memory Tokio channels to N waiting agents at zero marginal token cost.

Native sub-graph isolation for supervisor-worker patterns. Agents commit atomic JSON deltas with Optimistic Concurrency Control (OCC), avoiding monolithic snapshot bloat.

An immutable, versioned event ledger that rejects stale zombie writes and guarantees deterministic replay recovery in under 250ms upon container restart.

Single distroless Docker image with embedded RocksDB storage. 0.31ms marginal engine overhead (~2ms p50 durable commit latency), TLS native encryption, and zero external database dependencies for local and edge deployments.

Every execution step is journaled with full crash-durability. CellaFlow adds barely ~310µs in marginal engine overhead over raw loopback gRPC transport.

Zero-copy serialization of agent state payload into binary format in host memory.

Tonic gRPC socket transit, TLS frame packaging, and Bearer token interceptor validation.

Cognitive Graph validation, monotonic epoch fencing check, and embedded RocksDB WAL sync.

Single writer · Embedded RocksDB WAL fsync

When an agent commits state, the gRPC socket transport consumes 1.62ms. CellaFlow’s entire stateful persistence engine—including ledger serialization, idempotency validation, and WAL sync—takes only **310 microseconds**.

External databases introduce 10–25ms network hops and serialize agent requests under heavy connection pool contention. CellaFlow embeds RocksDB directly inside the 20MB daemon binary.

See how CellaFlow solves the hardest problems in agentic engineering, from runaway token bills to context collapse.

Prevent 100 autonomous agents from triggering API rate-limit bans or burning duplicate OpenAI tokens during parallel document synthesis.

Drop in `CellaflowSaver` to persist LangGraph state across spot recycles, Lambda timeouts, or container restarts with zero changes to your graph logic.

Guarantee that financial transactions, database writes, and external webhooks execute strictly at most once with deterministic idempotency keys.

Run 30-minute deep codebase refactoring agents safely with Non-Persistable Zones that protect token streams from checkpoint corruption.

See the immediate business case. Estimate how much LLM API budget you are throwing away on redundant steps due to infrastructure recycles and how Cellaflow's middleware stops the bleed.

Enable background log compaction (5x to 40x compression ratios) to optimize context inputs and save an extra ~45% in prompt token costs.

Directly thrown away on re-running completed steps.

Saved via asynchronous observation compaction.

Annual savings from Durable Replay + CCS Memory optimization.

Cellaflow maintains a dedicated compilation crate `cellaflow-proto` that compiles Protocol Buffer definitions (`proto/cellaflow/v1/`) on the `cargo build` phase using `tonic-build` in its build script. SDK clients and backend systems share exact runtime contracts without code replication.

Tonic gRPC server rejects unencrypted HTTP/2 immediately, securing all pipeline traffic.

Bearer Token metadata is validated at the gRPC interceptor layer before passing to the engine.

Locks active executions to the specific version they started on, preventing schema drift.

Bare loopback RPC is 1.62ms; CellaFlow's complete atomic WAL journaling and sequence fencing adds only 310µs.

```
syntax = "proto3";

package cellaflow.v1;

service WorkflowEngineService {
  // Starts a stateful session with registry version pinning
  rpc StartSession(StartSessionRequest) returns (StartSessionResponse);

  // Commits an execution step with lease validation & sequence guards
  rpc CommitStep(CommitStepRequest) returns (CommitStepResponse);

  // Recovers full Cognitive Graph history for transparent replay
  rpc GetGraph(GetGraphRequest) returns (GetGraphResponse);

  // Distributed idempotency cache & non-blocking lease heartbeats
  rpc CheckIdempotencyCache(CheckCacheRequest) returns (CheckCacheResponse);
  rpc RenewLease(RenewLeaseRequest) returns (RenewLeaseResponse);
  rpc ReleaseLease(ReleaseLeaseRequest) returns (ReleaseLeaseResponse);
}
```

The single-node engine is free, forever. One Docker command and a pip install is all you need.

Run the engine locally with Docker, install the Python SDK, and have your first durable workflow running in under 3 minutes.

Building something ambitious with multi-agent AI? Book a 30-min technical conversation. We'll review your architecture and help you integrate.

Prefer email? [hello@cellaflow.com](mailto:hello@cellaflow.com)
