{"slug": "beyond-the-ephemeral-loop-the-architecture-of-durable-ai-agents", "title": "Beyond the Ephemeral Loop: The Architecture of Durable AI Agents", "summary": "Earendil released @earendil-works/pi-durable at v1.1.0 on October 7, 2026, an experimental runtime that replaces the in-memory while-loop pattern of AI agents with an embedded SQLite storage engine and an operating system task scheduler. The package holds about 22,500 lines of TypeScript in src/, 72 test files, and a 263 KB design spec, and joins durable-execution offerings from Temporal, Restate, DBOS, Inngest, Cloudflare Workflows, and LangGraph checkpointers shipped across 2025 and 2026. Pi-Durable traces its lineage to Mario Zechner's @earendil-works/pi-coding-agent 1.0 terminal CLI, which relied on a human operator to restart crashed processes.", "body_md": "## On this page\n\n# Beyond the Ephemeral Loop: The Architecture of Durable AI Agents\n\nAI agents built on while loops lose state when processes crash. Durable agents require an embedded database and an operating system task scheduler.\n\nDevelopers build AI agents as a while loop around an LLM client.\n\nYou send a prompt, execute the requested tool, append the result to an in-memory array, and repeat. When a container restarts, a process crashes, or a network drops during a 45-second tool call, you lose the session. The user receives a severed socket. The database retains orphaned records. Recovery requires starting over.\n\n[Earendil](https://earendil.com) paired the release of [Pi 1.0](https://earendil.com/posts/pi-1-0/) with an experimental runtime: [`@earendil-works/pi-durable`](https://github.com/earendil-works/pi/tree/main/packages/durable). At [v1.1.0](https://github.com/earendil-works/pi/blob/main/packages/durable/CHANGELOG.md) (released October 7, 2026), the package holds about 22,500 lines of TypeScript in `src/`, 72 test files, and a 263 KB design spec.\n\nPi-Durable replaces the memory loop with an embedded storage engine and an operating system task scheduler.\n\nIt is not alone. In 2025 and 2026, nearly every agent platform shipped some form of durable execution: [Temporal](https://temporal.io), [Restate](https://restate.dev), [DBOS](https://docs.dbos.dev/architecture), [Inngest](https://www.inngest.com/docs/durable-execution/durable-agents), [Cloudflare Workflows](https://developers.cloudflare.com/workflows/), and [LangGraph checkpointers](https://docs.langchain.com/oss/python/langgraph/persistence). The defaults still bite. LangGraph’s `InMemorySaver` keeps every checkpoint in RAM, so a redeploy wipes them all. Most teams find out in production.\n\nWhat makes Pi-Durable worth studying is where it puts the durability: inside the process, next to the agent, in one SQLite file.\n\n## 1. From Terminal Loop to Durable Engine: The Evolution of Pi\n\nThe architecture of Pi-Durable reflects a concrete journey documented across two releases.\n\nPi originated as an interactive terminal CLI. Created by [Mario Zechner](https://mariozechner.at) at [Earendil](https://earendil.com), [`@earendil-works/pi-coding-agent`](https://github.com/earendil-works/pi/tree/main/packages/coding-agent) version 1.0 started as an antidote to bloated developer environments. As detailed in [Pi Agent: The Minimal Harness That Became My Multi-Tool Glue](https://kondasamy.com/blog/2026/pi-agent-multi-tool-workflow/), the design embraced radical subtraction: no native subagents, no plan mode, no permission prompts, and a prompt under 1,000 tokens.\n\nThat original runtime ran inside a terminal interface powered by [`@earendil-works/pi-tui`](https://github.com/earendil-works/pi/tree/main/packages/tui) with four built-in tools (`read`, `write`, `edit`, `bash`), appending messages to a [JSONL session tree](https://jsonlines.org/). If the process died, a human operator restarted the command-line interface and re-prompted the model. The operator acted as the crash recovery system.\n\n[`@earendil-works/pi-durable`](https://github.com/earendil-works/pi/tree/main/packages/durable) addresses autonomous server workloads. When an agent runs background triage, executes multi-step deployments, or supports team channels, you cannot rely on a human watching the terminal to restart broken processes.\n\nThis transition did not discard Pi’s minimalist foundation. As shown in [Harness Engineering: Stop Blaming the Model, Fix the Environment](https://kondasamy.com/blog/2026/harness-engineering-reliable-ai-agents/), reliable agent behavior requires engineering the operating context rather than expanding prompts. Pi-Durable moved that boundary into an embedded runtime.\n\n## 2. The Single Mutation Line: Storage Precedes Visibility\n\nStandard agents publish before they persist.\n\nAn agent streams text over a WebSocket: *“I cancelled your subscription and credited your balance.”* The process runs out of memory before executing the backing Stripe call. The customer reads the promise, but the transaction never occurred, and the in-memory transcript disappeared.\n\nPi-Durable prevents this race condition through a **single mutation line**.\n\nIn Pi-Durable, state changes pass through atomic storage commits:\n\n- Transcript entries (`pi.user` ,`pi.assistant` ,`pi.tool-result` )\n- Task state transitions (`pending` ,`running` ,`waiting` ,`terminal` )\n- Document mutations tracked as [Chord](https://github.com/earendil-works/pi/tree/main/packages/chord) JSON deltas\n- Queued inputs in the conversation inbox\n\nThe harness commits state to disk before streaming tokens to users. If any storage call throws, the Session fails and closes itself: every later call and every pending wait rejects with `SessionFailed`. Nothing retries. A half-written session is worse than a stopped one.\n\nCloudflare reached the same rule from the other side. A [Durable Object’s output gate](https://blog.cloudflare.com/durable-objects-easy-fast-correct-choose-three/) holds every outgoing message and response until the storage writes before it are confirmed. If the write fails, the message never leaves. Pi-Durable applies that rule to tokens, tool output, and UI state.\n\n``` js\nimport { openNodeSqliteStorage } from \"@earendil-works/pi-durable/storage/sqlite/node\";\nimport { Harness } from \"@earendil-works/pi-durable\";\n\n// Open the session over SQLite WAL\nconst harness = await Harness.open(\n  await openNodeSqliteStorage(\"./agent.sqlite\"),\n  { models, registry },\n  context\n);\n\nconst root = await harness.root(context);\n\n// Submitting work returns a durable submission handle\nconst submission = await root.submit(\n  {\n    type: \"input\",\n    content: \"Refactor auth middleware to use JWT\",\n    requestId: \"job-1042\" // Exactly-once deduplication key\n  },\n  context\n);\n\n// If the worker crashes here, the next process reopens storage and resumes\nconst settled = await submission.wait(context);\n```\n\nAnchoring user output in a storage commit removes network race conditions. A client reconnecting after a socket drop resumes from the committed database state.\n\n### What “Committed” Means on Each Backend\n\n“Committed” is only as strong as the storage under it. Pi-Durable ships four backends, and they make different promises:\n\n| Backend | Import | Survives process crash | Survives power loss | \n|---|---|---|---|\n| Memory | `MemoryStorage` | No | No | \n| SQLite | `storage/sqlite/node` | Yes | Newest commit may be lost | \n| JSONL | `storage/jsonl/node` | Yes | Yes, with `{ fsync: true }` | \n| Durable Object | `storage/sqlite/cloudflare` | Yes | Yes (Cloudflare replicates writes) | \n\nThe SQLite backend runs in [WAL mode](https://www.sqlite.org/wal.html) with `synchronous = NORMAL`. That is a deliberate trade. In WAL mode, a commit appends pages to a log file instead of rewriting the database, so readers never block the writer. With `NORMAL`, SQLite skips the fsync on each commit and syncs only at WAL checkpoints (by default, every 1,000 pages). A `kill -9` loses nothing. A pulled power cord can lose the last commit.\n\nFor a coding agent on a laptop, that is the right call. For a payment agent, pick JSONL with `fsync: true` or a Durable Object. Cloudflare [introduced SQLite-backed Durable Objects](https://blog.cloudflare.com/sqlite-in-durable-objects/) in September 2024 and made them generally available in April 2025, with up to 10 GB per object and 30 days of point-in-time recovery. One Durable Object per conversation gives every agent its own database and its own single writer.\n\nOne more constraint: one process owns a storage at a time. There is no cross-process locking. Point two workers at the same SQLite file and nothing stops both schedulers from running the same tasks.\n\n## 3. Document State: Bases, AST Splice Deltas, and Fork Pruning\n\nAgent state cannot live in freeform transcript prose alone. Complex workflows demand structured data: task lists, review diffs, active files, and billing ledgers.\n\nPi-Durable uses [Chord](https://github.com/earendil-works/pi/tree/main/packages/chord) to store typed JSON documents next to transcripts. Writing full JSON snapshots on every turn exhausts disk space, while storing raw diffs slows down read queries over time.\n\nThe storage engine balances this trade-off by alternating between bases and operational deltas.\n\nStorage writes changes as two record types:\n\n- Bases: Full JSON copies written at creation, during migrations, or when `checkpointWhen()` returns true.\n- Deltas: Compact Chord operations (`Op[]` ), such as array splices (`[\"p\", [\"items\"], 1, 0, [\"fix build\"]]` ) and value replacements (`[“s”, [“text”], “Done.”]).\n\nThe engine provides two history modes:\n\n1. `history: \"latest\"` : Storage deletes older deltas when a new base commits. Used by`pi.live` and`pi.inbox` to bound disk usage.\n2. `history: \"rewindable\"` : Storage preserves historical bases and deltas.`snapshotAsOf(entryId)` calculates document values as of past commit points.\n\n``` js\nimport { defineDoc } from \"@earendil-works/pi-durable\";\n\ninterface TodoState {\n  items: string[];\n}\n\nconst TodosDoc = defineDoc<TodoState>({\n  kind: \"app.todos\",\n  version: 1,\n  scope: \"conversation\",\n  history: \"rewindable\",\n  fork: \"asOf\", // Child conversation starts with parent state at fork entry\n  initial: () => ({ items: [] }),\n  // Commit a full base whenever two or more deltas accumulate\n  checkpointWhen: (_value, _ops, info) => info.deltasSinceBase >= 2,\n});\n```\n\n## 4. The Effect Sandwich: Executing External Effects Without Locks\n\nExternal effects cannot run inside a storage commit.\n\nA commit holds the database write queue. Executing an HTTP request, a Stripe charge, or a remote deployment inside a transaction stalls concurrent writers. Pi-Durable isolates side effects into a three-phase state machine: the Effect Sandwich.\n\n1. Commit Intent: The `prepare` phase generates an idempotency key derived from the task ID (`pay-42` ) and commits it to disk.\n2. Perform External Effect: The `charge` phase executes the network call using the committed key. This runs outside the database transaction, keeping storage unlocked.\n3. Commit Outcome: When the external call succeeds, the phase writes the terminal outcome and appends a receipt entry in one commit.\n\n``` js\nimport { defineTask } from \"@earendil-works/pi-durable\";\n\nexport const PaymentTask = defineTask<{ card: string }, any, any>({\n  name: \"shop.payment\",\n  version: 1,\n  initial: () => ({ phase: \"prepare\" }),\n  phases: {\n    // 1. Commit Intent\n    prepare: async (task, runtime, context) => {\n      const key = `pay-${task.id}`;\n      await runtime.commit(() => ({\n        status: \"running\",\n        checkpoint: { phase: \"charge\", key },\n      }), context);\n    },\n\n    // 2. Perform External Effect (outside commit line)\n    charge: async (task, runtime, context) => {\n      const key = task.state.checkpoint.key;\n      const receipt = await payments.charge(task.input.card, key);\n\n      // 3. Commit Outcome with receipt entry\n      await runtime.commit(async (tx) => {\n        const entry = await tx.appendEntry(runtime.conversationId, {\n          kind: \"app.receipt\",\n          data: { receiptId: receipt.id },\n        });\n        return {\n          status: \"terminal\",\n          outcome: { status: \"completed\", result: { entryId: entry.id } },\n        };\n      }, context);\n    },\n  },\n  abort: async (task, runtime, context) => {\n    await payments.refund(task.state.checkpoint.key);\n    await runtime.commit(() => ({\n      status: \"terminal\",\n      outcome: { status: \"aborted\", reason: \"user\" },\n    }), context);\n  },\n});\n```\n\n### Crash Recovery Taxonomy\n\nProcess crashes divide into three recovery states:\n\n| Crash Point | State on Reopening Storage | Recovery Action | \n|---|---|---|\n| **Before Step 1** | Checkpoint is `prepare` | Reruns `prepare` . No network call occurred. | \n| **Between Step 1 and Step 3** | Checkpoint is `charge` with`key: \"pay-42\"` | Reruns `charge` using the recorded key. The[Stripe Idempotency API](https://stripe.com/docs/api/idempotent_requests) recognizes the key and rejects duplicate charges. | \n| **After Step 3** | State is `terminal` with receipt | Skips execution. Work is complete. | \n\nCreating the idempotency key during the intent phase ties remote deduplication to the local database checkpoint. Generating the key inside the effect phase produces a new key on restart, breaking idempotency.\n\n**The key has a shelf life.** Stripe keeps idempotency keys for at least 24 hours, then prunes them. A task that crashes in `charge` and resumes three days later sends a key Stripe no longer remembers, and the charge runs again. Durable state on your side does not extend the provider’s memory. For effects that can sit for days, query the provider for an existing charge by your own reference before you retry.\n\n## 5. Tool Replay Policies: Safe vs. Unsafe Recovery\n\nProcess crashes during tool calls break standard agents.\n\nRerunning an active tool on restart triggers duplicate side effects, like double charges or repeat git commits. Dropping the active tool leaves the conversation broken. Pi-Durable requires every tool to declare a replay policy:\n\n``` js\nimport { Type } from \"@earendil-works/pi-ai\";\nimport { defineTool } from \"@earendil-works/pi-durable\";\n\nconst searchDocs = defineTool({\n  name: \"search_docs\",\n  description: \"Search internal knowledge base\",\n  parameters: Type.Object({ query: Type.String() }),\n  replay: \"safe\", // Safe to rerun on recovery\n  execute: async (args) => {\n    return { content: [{ type: \"text\", text: await index.query(args.query) }] };\n  },\n});\n\nconst executePayout = defineTool({\n  name: \"execute_payout\",\n  description: \"Release funds to customer\",\n  parameters: Type.Object({ amount: Type.Number(), accountId: Type.String() }),\n  replay: \"unsafe\", // Do not rerun after a crash\n  execute: async (args) => {\n    const receipt = await stripe.transfers.create({\n      amount: args.amount,\n      destination: args.accountId,\n    });\n    return { content: [{ type: \"text\", text: receipt.id }] };\n  },\n});\n```\n\nBefore `execute()` runs, the harness commits an execution intent checkpoint to disk:\n\n```\n{\n  \"phase\": \"execute\",\n  \"arguments\": { \"amount\": 250, \"accountId\": \"acct_8821\" },\n  \"replay\": \"unsafe\"\n}\n```\n\nWhen a crash occurs during tool execution:\n\n- `replay: \"safe\"` : The scheduler reruns the tool function on restart.\n- `replay: \"unsafe\"` : The scheduler refuses to rerun the tool. It synthesizes an`interrupted` result, packages whatever partial output was committed to`pi.live` before the crash, and hands control back to the model.\n\n### Testing Crash Recovery Under SIGKILL\n\nConsider a test scenario where a `kill -9` signal halts a process running two concurrent tool calls in one [SQLite](https://www.sqlite.org/wal.html) database:\n\n```\n// The SQLite row left by the killed process\n{\n  \"id\": 24,\n  \"kind\": \"pi.tool\",\n  \"owner\": 18,\n  \"state\": {\n    \"status\": \"running\",\n    \"checkpoint\": {\n      \"phase\": \"execute\",\n      \"arguments\": { \"target\": \"weekly\" },\n      \"replay\": \"safe\"\n    }\n  }\n}\n```\n\nWhen a second process opens the database:\n\n1. `Harness.open()` runs a reconciliation commit: all tasks marked`status: \"running\"` return to`status: \"pending\"` .\n2. The engine skips the unsafe `deploy` tool and appends an error entry:`\"[error] Tool deploy was interrupted and may have partially run\"` .\n3. The engine reruns the safe `fetch_report` tool from scratch.\n\nThe model receives the partial output, reads the interruption message, and decides how to proceed.\n\nThis is the key design choice: **the model, not the runtime, decides how to recover from an unsafe interruption.** The runtime knows a deploy may have half-run. Only the model, with the transcript in front of it, can decide whether to check status, roll back, or ask the user.\n\nUnsafe tools also shape what happens below them. A tool can call other tools as nested calls, each its own `pi.tool` task. When the calling tool is not replay-safe, its unfinished nested calls and child tasks get `abandonOnRestart: true` by default. On the next start, the scheduler aborts them, with everything they own, before any of them runs again. You never get an orphaned child finishing a job its parent will never collect.\n\n## 6. The Replay Paradox: Why Workflow Engines Clash With Agents\n\nEngineers evaluating durable execution often consider [Temporal](https://temporal.io).\n\nTemporal manages distributed microservices well. Running interactive, multi-turn agents on Temporal introduces architectural friction.\n\nThe two systems achieve durability through contrasting models:\n\n| Architectural Dimension | Temporal Workflow Engine | Pi-Durable Harness | \n|---|---|---|\n| **Durability Technique** | **Event Sourcing:** Re-runs workflow code from line 1, mocking completed activities. | **Checkpoint State Machine:** Never replays past code; loads the latest committed task state. | \n| **Code Determinism** | **Mandatory.** Calling`Date.now()` , random IDs, or changing an`if` block breaks history replay. | **Not required.** Code can change between turns. Only inputs, checkpoints, and outputs persist. | \n| **Code Hot-Reloading** | Complex. Requires patch markers ( `patched()` in TypeScript,`getVersion()` in Java and Go) or Worker Versioning. | **Native.** Swap the extension definition in the registry; the next phase runs the new code. | \n| **Token Streaming** | Inefficient. Pushing 500 token chunks through Temporal history hits payload and event limits. | **Throttled WAL commits.** Flushes token and tool stdout chunks every 100ms into a live document. | \n| **Runtime Infrastructure** | External cluster (Temporal Server + Postgres/Cassandra + gRPC workers). | **Embedded engine.** Runs inside Node, Bun, or[Cloudflare Durable Objects](https://developers.cloudflare.com/durable-objects/) over[SQLite WAL](https://www.sqlite.org/wal.html) . | \n\nTemporal enforces deterministic execution. If an agent executes tool calls for 40 minutes and the worker restarts, Temporal replays the workflow function from line 1. Every code branch must match the recorded event history.\n\nUpdating a prompt template, editing a system instruction, or modifying a tool definition during execution causes the replayed code to diverge from history. Temporal aborts with a non-deterministic workflow error.\n\nPi-Durable stores explicit checkpoints instead of replaying code. On restart, the engine loads the active checkpoint from SQLite, resolves the tool registry, and runs the next phase. Past code does not run again.\n\n### The Limits That Hit Agents First\n\nTemporal’s [hard limits](https://docs.temporal.io/cloud/limits) were set for business workflows, not chatty LLM loops:\n\n| Limit | Value | What it means for an agent | \n|---|---|---|\n| Payload size | 2 MB per input or result | A long tool output or a file read must be offloaded | \n| Event history | 51,200 events or 50 MB | Warnings start at 10,240 events or 10 MB | \n| Transaction / gRPC message | 4 MB | Caps a batch of tool results in one step | \n\nTemporal’s answer for big payloads is the claim-check pattern: write the blob to S3 and keep only a reference in history.\n\nRun the numbers on streaming. A 4,000-token answer streamed over 40 seconds, recorded every 100 ms, is 400 updates. Store each as an event and you hit the 51,200-event ceiling after 128 answers. A busy support conversation gets there in a day. So in practice teams stream tokens through a side channel (Redis, a WebSocket hub) and store only the final message in Temporal, which brings back the exact “user saw it, storage did not” gap from section 2.\n\nPi-Durable avoids the problem by not keeping history for live state. Partial output goes into `pi.live`, a document with `history: \"latest\"`. The harness flushes it at most every 100 ms with one commit in flight, and older deltas are deleted when a new base lands. The transcript gets one `pi.assistant` entry per response.\n\n### Where the Other Engines Sit\n\nDurable execution engines split into three camps by what they store:\n\n| Model | Engines | What a restart runs again | Code-change cost | \n|---|---|---|---|\n| Event sourcing | Temporal, Restate | Workflow code from the top, with recorded results fed back | Must stay deterministic; patch every change | \n| Step journaling | DBOS, Inngest, Cloudflare Workflows | Function from the top; finished steps return saved output | Step order must stay stable, or version the workflow | \n| State checkpoints | LangGraph, Pi-Durable | Nothing that finished; resume at the saved checkpoint | Code can change between steps | \n\n[DBOS](https://docs.dbos.dev/architecture) is the closest relative in spirit: a library, not a cluster, that writes each step’s output to Postgres and skips finished steps on recovery. But it still re-enters the workflow function from the top, so step order must match. LangGraph checkpoints graph state after every super-step, which is close to Pi-Durable’s model. The difference is scope: Pi-Durable also commits each tool call’s intent before it runs, owns the scheduler, and ties the transcript, documents, and tasks to one commit line.\n\nPi-Durable sits at the far end. It stores the state machine, not the function call stack. That is why live code swaps work, and also why you must write every phase so it can start cold from its checkpoint.\n\n## 7. The Step Check Engine and Live Code Handover\n\nReloading extension code during live runs requires isolated execution phases.\n\nIn Pi-Durable, each in-memory task run is an invocation. Invocations execute sequential phases. Between phases, the harness evaluates a check called a step.\n\nThe step executes as a commit on the Session write queue, verifying all changes committed by the prior phase before scheduling the next.\n\n### Preventing Infinite Loops With Step Faults\n\nIf a phase returns with the same checkpoint it started with, it loops on the same inputs.\n\nThe step compares old and new checkpoints as JSON, ignoring key order. When checkpoints match, the harness marks the task faulted:\n`\"Task phase returned without durable progress\"`.\n\nCommitting entries or documents does not satisfy this check. Only changing the checkpoint advances task state.\n\n### Live Code Handover\n\nAt each phase boundary, the step inspects the in-memory registry:\n\n- Updating an extension file on disk triggers `registry.install(NewExtension)` .\n- The step detects the updated definition for that task kind.\n- The task hands over: the engine returns it to `pending` with its current checkpoint, and the next scheduler pass executes the phase using the new definition.\n- Predecessor and successor execution handlers never overlap.\n\n### Handling Missing Extensions via the Blocked State\n\nUninstalling an extension or deploying a task version without a migration script pauses execution:\n\n- The harness does not terminate tasks when code is missing.\n- The task enters a `blocked` runtime state (`missing_task` ,`task_too_old` , or`migration_failed` ).\n- Storage retains the task in a pending state.\n- When you register the matching extension, the scheduler detects the definition and resumes execution.\n\n## 8. Hierarchical Concurrency: Bottom-Up Abort and Fail-Fast\n\nMost multi-agent implementations dispatch subagents as detached background async promises. If a user presses `Esc` or a payment fails, the parent halts, but child subagents continue burning API credits in the background.\n\nPi-Durable enforces hierarchical task ownership:\n\n``` js\nimport { defineTask, type TaskId } from \"@earendil-works/pi-durable\";\n\ntype CheckoutState =\n  | { phase: \"pay\" }\n  | { phase: \"decide\"; payments: TaskId<string>[] };\n\nconst Checkout = defineTask<{ cards: string[] }, CheckoutState, string>({\n  name: \"shop.checkout\",\n  version: 1,\n  initial: () => ({ phase: \"pay\" }),\n  phases: {\n    pay: async (task, runtime, context) => {\n      await runtime.commit(async (tx) => {\n        const payments: TaskId<string>[] = [];\n        for (const card of task.input.cards) {\n          payments.push(\n            await tx.createTask(Payment, { card }, {\n              ownership: { kind: \"task\", taskId: task.id }, // Child task\n            })\n          );\n        }\n        // Stop running code until all payments finish.\n        // The first failure aborts sibling payments.\n        return {\n          status: \"waiting\",\n          checkpoint: { phase: \"decide\", payments },\n          on: payments,\n          policy: \"failFast\",\n        };\n      }, context);\n    },\n    decide: async (task, runtime, context) => {\n      const outcomes = await runtime.outcomes(task.state.checkpoint.payments, context);\n      const paid = outcomes.every((outcome) => outcome.status === \"completed\");\n      await runtime.commit(() => ({\n        status: \"terminal\",\n        outcome: paid\n          ? { status: \"completed\", result: \"Order placed.\" }\n          : { status: \"failed\", error: { message: \"Payment failed.\" } },\n      }), context);\n    },\n  },\n  abort: async (task, runtime, context) => {\n    // Runs after child payments finish aborting and refunding\n    await runtime.commit(() => ({\n      status: \"terminal\",\n      outcome: { status: \"aborted\" },\n    }), context);\n  },\n});\n```\n\nTwo rules govern the task tree:\n\n1. `failFast` vs.`allSettled` : In`policy: \"failFast\"` , a failure in child task #2 marks siblings #1, #3, and #4 for cancellation.\n2. Bottom-Up Abort Order: When you abort a task or conversation, the cancel signal cascades to leaf tasks first. A parent task’s `abort()` handler runs only after child tasks reach a terminal state.\n\nCascading cancellations down to leaf tasks releases child holds, such as room bookings or card authorizations, before the parent marks the workflow failed.\n\nDeep trees used to be slow. The first scheduler visited every live task on each pass. The current one keeps in-memory ownership indexes and starts abort cascades only from tasks with cancellation intent. The [changelog](https://github.com/earendil-works/pi/blob/main/packages/durable/CHANGELOG.md) reports the gains:\n\n| Workload | Before | After | \n|---|---|---|\n| Chain of 500 owned tasks settling | 12 s | 50 ms | \n| Chain of 20,000 owned tasks settling | n/a | 1.5 s | \n| One tool making 8,000 parallel nested calls | 18.8 s | 2.7 s | \n| Tasks kept in memory after 10,000 ended subagent calls | 10,000 | 0 | \n\nThe last row matters for long sessions. A parent that spawns subagents all day no longer grows its memory with every ended child.\n\n## 9. The Semantic Inbox: Steers, Follow-Ups, and Passive Writes\n\nUsers do not wait for an agent to complete a 60-second tool turn before sending input. They interrupt.\n\nStandard architectures drop incoming messages or cancel running turns when users type. Pi-Durable routes incoming messages through a structured inbox (`pi.inbox`) supporting three operations:\n\n```\n// 1. STEER: Injected after the active tool round finishes.\n// Joins the running work to redirect the current turn.\nawait root.submit(\n  {\n    type: \"input\",\n    content: \"Don't use Postgres; use SQLite instead\",\n    whenBusy: \"steer\",\n  },\n  context\n);\n\n// 2. FOLLOW-UP: Queued until the agent finishes its current response.\n// Starts the next turn once idle.\nawait root.submit(\n  {\n    type: \"input\",\n    content: \"Now write the unit tests\",\n    whenBusy: \"followUp\",\n  },\n  context\n);\n\n// 3. WRITE: Appends an entry to the transcript without invoking the LLM.\n// Used for audit markers, file drop events, or CRM status sync.\nawait root.submit(\n  {\n    type: \"write\",\n    entry: { kind: \"app.audit\", data: { userTier: \"enterprise\" } },\n  },\n  context\n);\n```\n\nThere is a fourth mode: `whenBusy: \"reject\"` throws `ConversationBusy` instead of queueing, for callers like cron jobs that should never stack work. By default the inbox places one steer or follow-up per turn; setting `steeringMode: \"all\"` or `followUpMode: \"all\"` places every queued item at once. If a run fails, queued items stay in the inbox until the next submission places them, oldest first.\n\nWhen a tool finishes, the harness inspects the inbox before scheduling the next generation task. If a steer is pending, the engine injects the input into the ongoing run. The model receives the correction alongside the latest tool output, updating its trajectory without discarding prior execution.\n\n## 10. Two-Tier Compaction: Active Head vs. Permanent Audit\n\nEnterprise agents manage competing requirements: models have finite context windows, but compliance regulations (SOC 2, HIPAA, FINRA) require complete audit retention.\n\nPi-Durable decouples the model context window from the underlying storage transcript.\n\n1. Dual Compaction Thresholds:\n  - `backgroundTokens` (32,768 tokens below limit): A background compaction task summarizes older turns while user interactions proceed.\n  - `reserveTokens` (16,384 tokens below limit): If background summarization has not completed and the context window nears exhaustion, the next turn pauses until the summary settles.\n2. Head Pointers (`head: EntryId` ): Compaction writes a`pi.compaction` entry whose`head` references the oldest retained entry.\n3. Audit Preservation: Context reduction supplies entries from `head` forward to the model, warm with[Anthropic Prompt Caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) . The underlying database retains every raw entry from entry #1. Historical turns remain accessible through cursor queries (`tx.scanEntries()` ).\n4. Recent Context Stays Verbatim: `keepRecentTokens` (default 20,000) sets roughly how much of the newest context the summary leaves untouched.\n5. Overflow Recovery: When a provider rejects a request because the context is too long, generation compacts and retries once.\n6. Stale Summaries: A background summary that would cut before the start of the current context settles as `stale` . When several compactions are in flight, the furthest cut wins.\n\nThe cache detail is easy to miss. Each conversation stores a UUIDv7 in `pi.provider` and sends it as the provider `sessionId`. It survives reopen, retries, compaction, and model changes, so a restarted worker lands on the same prompt cache. With Anthropic, a cache read costs 10% of the base input price. A worker that crashes and resumes with a fresh session ID pays full price to rebuild a 150,000-token prefix; one that keeps the ID pays a tenth.\n\nFor a hard break, `reset()` starts a new context. It can take a handoff note (`\"We were fixing the flaky login test. Continue.\"`), and a tool can request one with `control: { handoff: \"...\" }`. The model stops seeing older entries; storage keeps them.\n\n## 11. Decision Framework\n\nChoose Pi-Durable when building:\n\n- Customer agents where process restarts must preserve conversational state\n- Terminal tools or edge services requiring an embedded footprint\n- Interactive systems needing live token streaming and mid-turn steering\n- Workloads targeting single-tenant SQLite databases or [Cloudflare Durable Objects](https://developers.cloudflare.com/durable-objects/)\n\nChoose [DBOS](https://docs.dbos.dev/) or [Inngest](https://www.inngest.com/) when building:\n\n- Agents that already live next to a Postgres database or a serverless function host\n- Step pipelines where step order rarely changes and you want each step’s output queryable as rows\n\nChoose [Temporal](https://temporal.io) when building:\n\n- Multi-service orchestrations spanning distributed databases and external teams\n- Asynchronous workflows that pause for weeks awaiting human approval\n- Batch billing, data synchronization, and server provisioning pipelines\n\n## The Bottom Line\n\nProduction AI agents demand systems engineering, not prompt tweaks.\n\nIn-memory loops fail under production traffic. Reliable agent runtimes require:\n\n- An atomic commit line that persists state before publishing output\n- An effect sandwich pattern isolating network calls from storage locks\n- Intent checkpoints paired with explicit tool replay rules\n- Hierarchical task trees with bottom-up abort cascades\n- A step verification engine that detects stalled loops and handles live code swaps\n- A semantic inbox routing steers, follow-ups, and background writes\n- Compaction that trims the active context window while retaining audit history\n\nPi-Durable demonstrates that reliable execution does not require a distributed server cluster. You can host crash-resilient, multi-conversation agents within an embedded library of about 22,500 lines on [SQLite](https://www.sqlite.org/).\n\n*Building durable agent architectures or wrestling with state management in production LLM apps? I would love to hear what patterns you are using. Reach out on [LinkedIn](https://www.linkedin.com/in/kondasamy/).*", "url": "https://wpnews.pro/news/beyond-the-ephemeral-loop-the-architecture-of-durable-ai-agents", "canonical_source": "https://kondasamy.com/blog/2026/pi-durable-architecture/", "published_at": "2026-10-11 00:00:00+00:00", "updated_at": "2026-10-11 06:22:03.856287+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "developer-tools", "artificial-intelligence", "mlops"], "entities": ["Earendil", "Pi-Durable", "@earendil-works/pi-durable", "Mario Zechner", "@earendil-works/pi-coding-agent", "Temporal", "Restate", "LangGraph"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/beyond-the-ephemeral-loop-the-architecture-of-durable-ai-agents", "markdown": "https://wpnews.pro/news/beyond-the-ephemeral-loop-the-architecture-of-durable-ai-agents.md", "text": "https://wpnews.pro/news/beyond-the-ephemeral-loop-the-architecture-of-durable-ai-agents.txt", "jsonld": "https://wpnews.pro/news/beyond-the-ephemeral-loop-the-architecture-of-durable-ai-agents.jsonld"}}