Pi Durable Earendil and the Pi community shipped Pi 1.0 alongside Pi Durable, an experimental package for building long-running, durable agents that run anywhere a JavaScript runtime exists. Pi Durable ships memory, SQLite, and JSONL storage backends plus a conformance suite and benchmarks, with SQLite and JSONL code using no Node APIs so it can run on Bun or inside a Cloudflare Durable Object. The full source, excluding tests, is about 15,000 lines — roughly 150,000 tokens with GPT and 250,000 with Claude — and the storage backends alone account for 3,000 lines. Pi Durable Today Earendil and the Pi community shipped Pi 1.0 https://earendil.com/posts/pi-1-0/ . This reflects our belief that after countless hours of hardening, maintenance, and active development, Pi is now a solid foundation on which to build. Pi also continues to evolve. Together with Pi 1.0, we are shipping an experimental new package called Pi Durable. Pi Durable was built specifically for long-running, durable, and malleable agents that can run anywhere. We would like you to join in the fun and help us make it the best durable harness there is. Why Pi Durable? Pi the coding agent is built to run on your remote machine, inside a terminal, driven by one person. If the process dies, you look at what happened and tell it to continue. That is what Pi 1.0 focuses on and excels at, and that is not changing. At Earendil, we want to bring this technology to everyone, in whatever form fits their needs best. For that, we need a harness that runs anywhere, can be reached from different surfaces, supports infinitely long conversations, survives catastrophic internal and external failures, and lets multiple humans steer the same agents. Pi Durable is that harness. It does not replace the Pi coding agent. It is a framework for building any agentic application, coding agents included. It shares not only code with the Pi coding agent, like pi-ai, but also its principles: minimalism and malleability. It also lets us explore designs in this space without disrupting Pi the coding agent. Lessons we learn building agentic applications on Pi Durable will flow back into Pi the coding agent as they prove themselves valuable. What is a harness? Everybody has their own definition of a harness. We wrote about this previously https://earendil.com/posts/what-is-a-harness/ , but let us reintroduce the concept of the harness for Pi Durable. A harness is storage plus the machinery needed to run one or more conversations with large language models in parallel. It provides the tools those models call, and the execution environments the tools run in. A conversation is an interaction between you and an agent, recorded as a transcript. The agent is the large language model together with its settings, like the thinking level, and the tools it can call. Tools do their work through an execution environment, which can be your laptop, a remote VM, or an in-memory sandbox. Which tools and which execution environment an agent gets is up to each conversation. Everything the harness runs, from calling the model to executing a tool, is a task. Like everything in Pi, Pi Durable is built so your agent can understand it. The entire source code, without tests, is about 15,000 lines, which comes out to about 150,000 tokens with GPT and about 250,000 with Claude. That's the worst case. To build on Pi Durable, your agent rarely needs all of it; the storage backends alone are 3,000 lines it can usually skip. Now let us give you a little tour of Pi Durable, to illustrate what we built and why we built it. Long runs anywhere We want agents to run for a long time and to be able to run anywhere, where anywhere currently means anywhere there is a JavaScript runtime. In Pi Durable, a harness opens over a storage backend. Pi Durable ships memory, SQLite, and JSONL storage, plus a conformance suite and benchmarks for your own backend. The SQLite and JSONL storage code uses no Node APIs, so with a small adapter it runs on Bun or inside a Cloudflare Durable Object. The storage interface is small and easy to implement on top of whatever you have, like a key-value store or Postgres. One process owns a storage at a time, and other clients attach to that process. On SQLite, the harness only keeps the working set in memory: the active transcripts, live tasks, and pending submissions. Everything else stays on disk until it is needed. Active transcripts are naturally bounded by the model's context window, because compaction summarizes older messages before they overflow it. So even a conversation with tens of thousands of messages fits snugly into memory. Tools that need files or a shell get them from an execution environment. Pi Durable ships a Node execution environment, which gives tools access to your local files. Like storage, the execution environment interface is small and easy to implement, so you can also expose remote execution environments to your tools. That allows the harness to run on one machine while its tools run on another. Your env function builds the environment for every tool call, from the conversation's working directory, so each conversation can run in a different place. js import { BACKGROUND CONTEXT } from "@earendil-works/chord/context"; import { createModels } from "@earendil-works/pi-ai/models"; import { openaiProvider } from "@earendil-works/pi-ai/providers/openai"; import { createRegistry, Harness } from "@earendil-works/pi-durable"; import { NodeExecutionEnv } from "@earendil-works/pi-durable/env/node"; import { openNodeSqliteStorage, } from "@earendil-works/pi-durable/storage/sqlite/node"; import { CodingTools } from "@earendil-works/pi-durable/tools"; const context = BACKGROUND CONTEXT; // every call takes a context for cancellation const models = createModels ; models.setProvider openaiProvider ; const registry = createRegistry ; registry.install CodingTools ; // read, write, edit, bash const env = { cwd }: { cwd?: string } = new NodeExecutionEnv { cwd: cwd ?? process.cwd } ; const harness = await Harness.open await openNodeSqliteStorage "./agent.sqlite" , { models, registry, env }, context, ; // The root conversation: created on first use, and the same one after every // restart. const root = await harness.root context, { agent: { model: { provider: "openai", modelId: "gpt-6.1-sol" }, cwd: "/work/repo", }, } ; Survives crashes We want an agent to survive its process dying, whether the laptop sleeps, the container is redeployed, or the machine runs out of memory, and to pick up where it left off. In Pi Durable, every step of a run is a task that stores a checkpoint before it moves on. If the process dies, a new process opens the same storage, finds the unfinished tasks, and continues each one from its last checkpoint. A model request that was cut off is sent again; the partial answer stays in the transcript, marked as aborted. A tool call that was cut off reruns if it is safe to; otherwise the model is told it was interrupted. Pi Durable has no built-in subagents, but they take a few lines of code to build, as the triage tool below shows. A subagent runs in a conversation of its own, so it continues from where it left off too, and a subagent tool that is safe to rerun finds its subagent again and waits for its answer. Queued messages are still queued. A requestId makes a submission exactly-once, so a client that retries after a crash gets the original submission back instead of asking twice. js const job = { type: "input", content: "Fix the flaky login test", requestId: "job-42", } as const; await root.submit job, context ; // The process dies here, in the middle of a tool call. // A new process opens the same storage. const harness = await Harness.open await openNodeSqliteStorage "./agent.sqlite" , { models, registry, env }, context, ; harness.resume ; // continue the interrupted run const root = await harness.root context ; // the same submission, answered const settled = await await root.submit job, context .wait context ; Many conversations at once We want one harness to run many conversations at the same time, without one blocking another. In Pi Durable, one harness runs as many conversations as you need, concurrently, all with the same guarantees. A conversation starts fresh or forks another one at any point in its transcript, and sees the parent's history up to that point without copying it. Think of a Slack channel where your agent answers mentions from anyone. Then somebody opens a thread. The channel can be one conversation, and the thread a fork of it at the message the thread replies to. Both run at the same time and neither blocks the other. js const channel = await harness.root context ; const question = await channel.submit { type: "input", content: "@agent why did the deploy fail?" }, context, ; const answered = await question.wait context ; // Someone replies to the agent's answer in a thread. Every conversation names // its owner, which decides what an abort reaches more on that under Tasks . // The thread has none. const thread = await channel.fork answered.answer , { ownership: { kind: "ownerless" } }, context, ; // Both conversations work at the same time. const inThread = await thread.submit { type: "input", content: "@agent can we roll it back?" }, context, ; const inChannel = await channel.submit { type: "input", content: "@agent who is on call today?" }, context, ; await Promise.all inThread.wait context , inChannel.wait context ; Each conversation also stores its own agent: the model, the thinking level, the selected extensions and which of their tools are active, extra instructions, and the working directory in its execution environment. A reviewer next to the main agent can use a cheaper model, read-only tools, and its own checkout. Extensions We want everything an agent can do to be pluggable, and every plugged-in piece to take part in durability. In Pi Durable, an extension is a named bundle of system prompt sections, tools, hooks, and tasks. The application installs extensions in a registry. Each conversation selects which extensions and tools it uses, and stores only their names. System prompt sections The system prompt is rebuilt from the sections of the conversation's extensions before every request, so a changed section is picked up by the next request. Pi Durable records what changed in the transcript, at the position where it changed, so a restart or a fork sees exactly what the model saw. On models that support system prompt and tool changes in the middle of a conversation, only the change is sent, so the prompt cache stays valid. js import { defineExtension, section } from "@earendil-works/pi-durable"; const ProjectContext = defineExtension { name: "project-context", sections: // Read from the conversation's execution environment. The files can be // loaded and watched in the background; every request renders the // latest state. section "agents md", input = agentsMd.latest input.env , section "skills", input = skills.latest input.env , , } ; Tools Every tool call runs as its own durable task, and its intent is stored before it runs. After a crash, a tool reruns only if it says that is safe. Otherwise the model is told the call was interrupted, with the output stored so far, and decides what to do. Each conversation can also get its own set of tools, like the Slack thread from earlier, which may search but not deploy. js import { Type } from "@earendil-works/pi-ai"; import { defineTool } from "@earendil-works/pi-durable"; const searchIssues = defineTool { name: "search issues", description: "Search the issue tracker", parameters: Type.Object { query: Type.String } , replay: "safe", // only reads, so a rerun after a crash is fine execute: async args, api = { // streamed to every client watching api.output searching for ${args.query}\n ; return { content: { type: "text", text: await tracker.search args.query } , }; }, } ; const deploy = defineTool { name: "deploy", description: "Deploy a version to production", parameters: Type.Object { version: Type.String } , // No replay: a deploy interrupted by a crash is reported to the model, // never repeated. execute: async args = { content: { type: "text", text: await ci.deploy args.version } , } , } ; registry.install defineExtension { name: "ops", tools: searchIssues, deploy } ; // The thread may search, but not deploy. await thread.configure { tools: { remove: deploy } }, context ; A tool gets the harness API for its call: it can commit entries and documents, start tasks and conversations, and talk to other conversations. That makes a subagent a few lines of code. A tool creates a conversation it owns, gives it a smaller model and its own instructions, and waits for its answer. The subagent is a conversation like any other, so it survives a crash, counts its own cost, and a UI can show it under the call. python import type { AssistantMessage } from "@earendil-works/pi-ai"; import { AssistantEntry, configure } from "@earendil-works/pi-durable"; const triage = defineTool { name: "triage", description: "Label an incoming issue as bug, feature, or question", parameters: Type.Object { issue: Type.String } , // a rerun after a crash finds the same subagent and the same submission replay: "safe", execute: async args, api, context = { const child = await api.commit async tx = { const existing = await tx.scanConversations { ownerTaskId: api.taskId }, 1 .items 0 ; if existing == undefined return existing.id; // Owned by this call, so aborting the call aborts the subagent. const created = await tx.createConversation { ownership: { kind: "task", taskId: api.taskId }, } ; // It starts as a copy of this conversation's agent. Make it a small // model without tools. await configure tx, created.id, { model: { provider: "openai", modelId: "gpt-6-luna" }, tools: , instructions: "Answer with one word: bug, feature, or question.", } ; return created.id; }, context ; // lets a UI show the subagent under the call await api.details { conversationId: child }, context ; const subagent = await api.conversation child, context ; const request = { type: "input", content: args.issue, requestId: triage:${api.taskId} , } as const; const settled = await await subagent .submit request, context .wait context ; // The answer is an entry in the subagent's transcript. Read it and take // its text. const entry = await api.commit tx = tx.entry AssistantEntry, settled.answer , context, ; const message = entry?.model?. 0 as AssistantMessage; const text = message.content .flatMap content = content.type === "text" ? content.text : .join "" ; return { content: { type: "text", text } }; }, } ; Extensions can also change other extensions' tools. A tool with the same name in a later extension replaces the earlier one, for example a bash that runs inside a Python virtualenv. A wrap decorates whichever tool won, wherever the wrapping extension is selected. js import { wrapTool } from "@earendil-works/pi-durable"; import { createBashTool } from "@earendil-works/pi-durable/tools"; // Times every bash call, whichever bash the conversation ends up with. const Timing = defineExtension { name: "timing", wraps: wrapTool createBashTool , bash = { ...bash, execute: async args, api, context = { const start = Date.now ; try { return await bash.execute args, api, context ; } finally { metrics.record "bash", Date.now - start ; } }, } , , } ; Hooks Hooks let extensions step into tasks, including the built-in tasks for generating a model response, invoking a tool, or performing compaction. They can rewrite a request before it goes to the model, block or rewrite a tool call, replace a result, keep a run going, or write a summary themselves. A hook can run again after a crash, so a hook that makes a decision stores it in a memo: a small value stored with the task, where the first write wins. js import { hook, ToolTask } from "@earendil-works/pi-durable"; const Approval = defineExtension { name: "approval", hooks: hook ToolTask, { beforeTool: async call, api, context = { if call.name == "deploy" return undefined; // After a restart, the hook finds the stored answer instead of // asking again. let approved = await api.memo