{"slug": "building-a-pi-extension-with-jev-agent-hooks-supervision-and-real-tests", "title": "Building a Pi Extension with Jev: Agent Hooks, Supervision, and Real Tests", "summary": "A developer built Jev Loop Control, a Pi coding-agent extension that uses a second model to assess an actor agent's decisions and evidence at key points in a session, aiming to catch cases where an agent silently drops a requirement while still reporting success. The extension hooks into Pi events such as message_end, tool_call, tool_result, context, and agent_before_settle, and the developer notes the control paths work in tests but the experiments do not yet establish improved coding outcomes. The writeup includes a sample extension that blocks Pi's write and edit tools from modifying a project's requirements.md.", "body_md": "My coding agent can edit files, run tests, and keep working without me. The harder question is whether it is still working toward the right thing.\n\nSuppose I ask for a configuration parser that **rejects unknown settings**. The agent decides to accept them instead. It can implement that decision perfectly, write tests for it, and announce success. The code works; the requirement was lost.\n\nI built [Jev Loop Control](https://github.com/krisitown/jev-loop-control) to explore whether a second model could catch problems like this at useful moments in a Pi session. Pi runs the tools, the actor writes the code, and Jev assesses selected decisions and evidence. Extension code decides what happens next.\n\nThe control paths work in tests. **The experiments do not yet establish that it makes coding outcomes better.** That distinction shaped both the implementation and the next tests.\n\nThis is the practical companion to my video: how the hooks fit together, a small extension you can run, how to try Jev Loop Control, and a short account of what we learned. The full methods, measurements, and limitations are in the [technical report on Zenodo](https://zenodo.org/records/23046412).\n\n[Pi](https://github.com/earendil-works/pi) is the coding-agent harness in this project. My actor was a local Qwen model running on an NVIDIA DGX Spark, but the extension uses the model already configured in Pi; it does not require that hardware.\n\nThe useful extension points sit around the loop: a completed assistant message, a tool about to execute, its result, the context for the next request, and the point where the agent is about to finish. You can add behavior at these boundaries without rewriting the agent.\n\n| Pi API or event | What I use it for | \n|---|---|\n| `message_end` | Capture the complete assistant proposal and its source. | \n| `tool_call` | Apply an existing hold before a dependent tool executes. | \n| `tool_result` | Record what actually happened, including errors and verification output. | \n| `context` | Add bounded policy guidance and selected memory to the next request. | \n| `registerTool` | Give the actor a structured `jev_checkpoint` tool. | \n| `agent_before_settle` | Check a completion claim and, when justified, request a bounded continuation. | \n| `agent_settled` | Write the final summary; this event is notification-only. | \n\n**The examples below target Pi 0.87.1**, the version used for integration validation. See the [Pi extension documentation](https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/extensions.md) for the current API. Tool calls from the same message may execute in parallel, so a checkpoint must be called and awaited **before** dependent work, not alongside it.\n\nBefore involving another model, make one rule work end to end. This example prevents Pi's built-in `write` and `edit` tools from changing a project's `requirements.md` file. It adds a `/contract-status` command so you can see what is protected.\n\nSave this as `protect-requirements.ts`:\n\n``` js\nimport { resolve } from \"node:path\";\nimport {\n  isToolCallEventType,\n  type ExtensionAPI,\n} from \"@earendil-works/pi-coding-agent\";\n\nexport default function (pi: ExtensionAPI) {\n  pi.registerCommand(\"contract-status\", {\n    description: \"Show the protected requirements file\",\n    handler: async (_args, ctx) => {\n      ctx.ui.notify(`Protected: ${resolve(ctx.cwd, \"requirements.md\")}`, \"info\");\n    },\n  });\n\n  pi.on(\"tool_call\", async (event, ctx) => {\n    if (\n      !isToolCallEventType(\"write\", event) &&\n      !isToolCallEventType(\"edit\", event)\n    ) return;\n\n    const target = resolve(ctx.cwd, event.input.path);\n    if (target !== resolve(ctx.cwd, \"requirements.md\")) return;\n\n    return {\n      block: true,\n      reason: \"Keep requirements.md unchanged. Implement the requirement in code.\",\n    };\n  });\n}\n```\n\nWith Node 22.19.0 or newer, install the tested Pi version if you do not already have it:\n\n```\nnpm install -g @earendil-works/pi-coding-agent@0.87.1\n```\n\nIn a scratch project, create `requirements.md` containing “Unknown settings must be rejected.” Then load the example:\n\n```\npi --extension ./protect-requirements.ts\n```\n\nUse your configured Pi model and run `/contract-status`. Ask it to read the file, then to use the `edit` tool to change “rejected” to “accepted.” If it attempts that edit, Pi should return the block reason before the edit executes. Edits to other files remain available.\n\nThe default export receives `ExtensionAPI`; it registers a command and an event handler. The handler narrows the tool type, checks its target, and returns Pi's supported blocking result. Pi loads TypeScript extensions directly, so this example needs no separate build step.\n\nThis is a teaching example, **not a filesystem sandbox**: it does not cover shell commands, other tools, or alternate paths through symlinks. It also cannot tell whether an implementation respects the requirement. That semantic question is where the Jev experiment begins.\n\nIn the real extension, a judgment starts with a specific event and a bounded evidence packet. The packet separates user requirements, observed tool results, actor claims, and hypotheses. A claim that “the tests passed” is different from the recorded output of a test command.\n\nThe sequence is:\n\nJev's output is input to a policy. It is not executable authority. A concern without a usable evidence anchor may produce no intervention.\n\nThe source separates these responsibilities: [`src/policies/`](https://github.com/krisitown/jev-loop-control/tree/v0.5.0-dev.1/src/policies) holds policy decisions, while [`src/policy-live.ts`](https://github.com/krisitown/jev-loop-control/blob/v0.5.0-dev.1/src/policy-live.ts) connects them to Pi. This makes it possible to test delivery with controlled judgments before asking whether hosted judgments are useful.\n\nThe earlier supervisor inspected proposals broadly. The redesign asks narrower questions at distinct boundaries:\n\n| Policy | Trigger | Useful question | \n|---|---|---|\n| **Context Governor** | Actor offers a source-linked fact; memory is selected before a request. | What should we retain or bring back into view? | \n| **Decision Gate** | Actor reports a consequential decision. | Does this choice conflict with the requirement or lack critical evidence? | \n| **Assumption Gate** | Actor reports an assumption it is about to rely on. | Can we carry it forward, verify it cheaply, or do we need user input? | \n| **Progress Governor** | Actor reports progress, or enough tool results accumulate. | Are we learning something, or repeating an unproductive approach? | \n| **Completion Gate** | Pi reaches its completion boundary. | Does current evidence support the claimed result? | \n\nDecision and assumption checks depend on the actor reporting them through `jev_checkpoint`. An unreported decision can still be missed. Progress and completion also have automatic triggers. These names describe this project's modules, not a standard taxonomy.\n\nFor the parser example, tests that accept valid settings do not establish that unknown settings are rejected. A completion check needs evidence for the requested behavior, tied to the code revision being claimed complete. An old passing test cannot automatically clear a later edit.\n\nThe video includes a **planted conflict**: the requirement says to reject unknown settings, while the proposed change accepts them. That demonstrates a control path. We also needed ordinary tasks, productive controls, and independent acceptance checks.\n\nThe [report](https://zenodo.org/records/23046412) separates the studies because they answer different questions. They span development stages; this is not one benchmark of a single frozen version:\n\n| Study | Observed result | What it supports | \n|---|---|---|\n| **13 final-source Pi smoke scenarios** | 10 of 12 Qwen cases passed their declared assertions. A separate scripted actor's dependent write was blocked. Two Qwen checks missed their original expectations and remain reported. | Controlled judgments can reach Pi and affect execution. This does not measure hosted-model usefulness. | \n| **Three-task paired pilot** | Both arms scored **14/15** acceptance checks. Jev logged**122 assessments and 76 interventions** , with greater recorded time and token use. | More intervention did not produce a higher score in this pilot. | \n| **Seeded policy calibration** | At exploratory thresholds of **0.50 probability / 0.10 margin** , only context selection met its gate: four repeated challenge observations and four controls. | A promising result on authored packets, not eight independent held-out tasks or proof of context savings. | \n| **Short decision challenge and control** | Plain Pi and a generic checklist completed both. Both Jev continuations timed out; three usable judgments produced no corrective effect. | Cleaner diagnostic prompts did not transfer reliably to these live continuations. | \n| **Adaptive sample-project campaign** | Mixed outcomes. The later B9 retry ended at **253/256** checks in both arms, with**30 Jev assessments and no applied interventions** . | Useful traces to investigate; changing versions, different starting histories, and incomplete runs prevent a causal comparison. | \n\nIn the three-task pilot, recorded input tokens were about **5.4×** higher with Jev; actor time was about **1.6×** higher. These are totals from a tiny exploratory study, not a forecast for other projects or proof that supervision caused every part of the overhead.\n\nThree lessons changed the design. **Zero interventions is ambiguous:** the actor may already have corrected itself, or the evaluator may have been unavailable or insufficiently grounded. **Successful execution is not successful verification:** a command can finish while its output reports failing tests. And **a better prompt on clean examples is not enough:** the question must still work with the noisy evidence a real run supplies.\n\nAll five policies are implemented. General coding benefit remains unproven. Context selection is the most promising lead from these tests, but we have not demonstrated that it reduces the actual actor prompt while preserving held-out acceptance.\n\nThe [source, README, and configuration example](https://github.com/krisitown/jev-loop-control/tree/v0.5.0-dev.1) are public. The version discussed here is **0.5.0-dev.1**, a development release, tested with Pi **0.87.1** and Node **22.19.0+**.\n\nInstall the pinned version:\n\n```\npi install git:github.com/krisitown/jev-loop-control@v0.5.0-dev.1\nexport AI_GATEWAY_API_KEY=\"your-vercel-ai-gateway-key\"\n```\n\nUse a Vercel AI Gateway key with access to `typesafe-ai/jev`. Start a new Pi process after exporting it. The extension does not automatically load `.env` files. Your actor can remain local; bounded assessment packets go to the configured Jev endpoint, by default Vercel AI Gateway's TypeSafe endpoint.\n\nFor a first inspection, create `jev-loop-control.config.json` in your test project:\n\n```\n{\n  \"mode\": \"observe\",\n  \"budget\": { \"maxRequests\": 20 },\n  \"limits\": {\n    \"maxAssessments\": 20,\n    \"maxInterventionsPerTask\": 4,\n    \"maxTerminalContinuations\": 1\n  }\n}\n```\n\nThen start `pi` and run:\n\n```\n/jev-status\n/jev-doctor\n/jev-trace\n```\n\n`observe` records judgments without applying guidance, blocks, or policy memory promotion. It still makes hosted calls. It also omits the enforcement-mode checkpoint instructions, so it is not a behaviorally identical rehearsal of every decision and assumption check.\n\n**Without a configuration file, the package defaults to `enforce`, all five policies enabled, and unlimited requests, assessments, and interventions.** The small limits above are deliberate choices for this walkthrough, not the package defaults. A limit being reached is not a successful verification.\n\nWhen you want to exercise supported effects, restart in enforcement mode:\n\n```\nJEV_LOOP_CONTROL_MODE=enforce pi\n```\n\nThe environment mode overrides the file; its budgets still apply. Give the agent a small task with a clear, independently testable contract. For example:\n\nImplement a settings parser that accepts `port` and `logLevel`, rejects every unknown key, and tests both behaviors. Before making a consequential design choice, call `jev_checkpoint` separately and await the result. Cite available evidence sources and report verification only after the relevant tool has run.\n\nThis is a normal task, not a guarantee of an intervention. A correctly proceeding actor is a useful control. Do not lower thresholds simply to manufacture a block.\n\nUse the trace directory reported by `/jev-trace` to inspect `events.jsonl` and, when enabled, the sanitized request/response artifacts. Follow **event → evidence → judgment → effect → delivery → subsequent behavior**. A recorded recommendation is different from a delivered block, and a delivered block is different from a better result. Project content in traces can still be sensitive even though credentials are scrubbed.\n\nI want to isolate context selection: start every arm from the same task and history, place an important fact early, and later require it. Add irrelevant and superseded facts so retaining everything cannot win by default.\n\nCompare plain history, simple deterministic selection or summaries, and Jev selection. Measure independent acceptance, the **actual context sent to the actor**, tokens, time, and harmful omissions. For decision checks, include a generic “check your approach” reminder as a baseline. If that reminder helps just as much, the reminder deserves the credit.\n\nThe goal is an agent I need to interrupt less. That has to show up in outcomes, not just a busy supervisor log.\n\n`10.5281/zenodo.23046412`. It has not been peer reviewed.\nThe full campaign traces are not all bundled with the extension repository. The report identifies the studies and their limitations; the repository provides the implementation and its tests.\n\nIf you have tried supervising a coding agent, what did you measure: a missed bug caught, fewer repair rounds, less context, or mostly another source of latency? A small, reproducible comparison against a simpler alternative would be especially useful.", "url": "https://wpnews.pro/news/building-a-pi-extension-with-jev-agent-hooks-supervision-and-real-tests", "canonical_source": "https://dev.to/kstoyanovai/building-a-pi-extension-with-jev-agent-hooks-supervision-and-real-tests-150a", "published_at": "2026-09-29 20:59:22+00:00", "updated_at": "2026-09-29 21:16:50.474924+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models", "ai-safety"], "entities": ["Pi", "Jev Loop Control", "Qwen", "NVIDIA DGX Spark", "Zenodo", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-a-pi-extension-with-jev-agent-hooks-supervision-and-real-tests", "markdown": "https://wpnews.pro/news/building-a-pi-extension-with-jev-agent-hooks-supervision-and-real-tests.md", "text": "https://wpnews.pro/news/building-a-pi-extension-with-jev-agent-hooks-supervision-and-real-tests.txt", "jsonld": "https://wpnews.pro/news/building-a-pi-extension-with-jev-agent-hooks-supervision-and-real-tests.jsonld"}}