Building a Pi Extension with Jev: Agent Hooks, Supervision, and Real Tests A developer built Jev Loop Control, a Pi coding-agent extension that uses a second model to assess an actor agent's decisions and evidence at key points in a session, aiming to catch cases where an agent silently drops a requirement while still reporting success. The extension hooks into Pi events such as message_end, tool_call, tool_result, context, and agent_before_settle, and the developer notes the control paths work in tests but the experiments do not yet establish improved coding outcomes. The writeup includes a sample extension that blocks Pi's write and edit tools from modifying a project's requirements.md. My coding agent can edit files, run tests, and keep working without me. The harder question is whether it is still working toward the right thing. Suppose I ask for a configuration parser that rejects unknown settings . The agent decides to accept them instead. It can implement that decision perfectly, write tests for it, and announce success. The code works; the requirement was lost. I built Jev Loop Control https://github.com/krisitown/jev-loop-control to explore whether a second model could catch problems like this at useful moments in a Pi session. Pi runs the tools, the actor writes the code, and Jev assesses selected decisions and evidence. Extension code decides what happens next. The control paths work in tests. The experiments do not yet establish that it makes coding outcomes better. That distinction shaped both the implementation and the next tests. This is the practical companion to my video: how the hooks fit together, a small extension you can run, how to try Jev Loop Control, and a short account of what we learned. The full methods, measurements, and limitations are in the technical report on Zenodo https://zenodo.org/records/23046412 . Pi https://github.com/earendil-works/pi is the coding-agent harness in this project. My actor was a local Qwen model running on an NVIDIA DGX Spark, but the extension uses the model already configured in Pi; it does not require that hardware. The useful extension points sit around the loop: a completed assistant message, a tool about to execute, its result, the context for the next request, and the point where the agent is about to finish. You can add behavior at these boundaries without rewriting the agent. | Pi API or event | What I use it for | |---|---| | message end | Capture the complete assistant proposal and its source. | | tool call | Apply an existing hold before a dependent tool executes. | | tool result | Record what actually happened, including errors and verification output. | | context | Add bounded policy guidance and selected memory to the next request. | | registerTool | Give the actor a structured jev checkpoint tool. | | agent before settle | Check a completion claim and, when justified, request a bounded continuation. | | agent settled | Write the final summary; this event is notification-only. | The examples below target Pi 0.87.1 , the version used for integration validation. See the Pi extension documentation https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/extensions.md for the current API. Tool calls from the same message may execute in parallel, so a checkpoint must be called and awaited before dependent work, not alongside it. Before involving another model, make one rule work end to end. This example prevents Pi's built-in write and edit tools from changing a project's requirements.md file. It adds a /contract-status command so you can see what is protected. Save this as protect-requirements.ts : js import { resolve } from "node:path"; import { isToolCallEventType, type ExtensionAPI, } from "@earendil-works/pi-coding-agent"; export default function pi: ExtensionAPI { pi.registerCommand "contract-status", { description: "Show the protected requirements file", handler: async args, ctx = { ctx.ui.notify Protected: ${resolve ctx.cwd, "requirements.md" } , "info" ; }, } ; pi.on "tool call", async event, ctx = { if isToolCallEventType "write", event && isToolCallEventType "edit", event return; const target = resolve ctx.cwd, event.input.path ; if target == resolve ctx.cwd, "requirements.md" return; return { block: true, reason: "Keep requirements.md unchanged. Implement the requirement in code.", }; } ; } With Node 22.19.0 or newer, install the tested Pi version if you do not already have it: npm install -g @earendil-works/pi-coding-agent@0.87.1 In a scratch project, create requirements.md containing “Unknown settings must be rejected.” Then load the example: pi --extension ./protect-requirements.ts Use your configured Pi model and run /contract-status . Ask it to read the file, then to use the edit tool to change “rejected” to “accepted.” If it attempts that edit, Pi should return the block reason before the edit executes. Edits to other files remain available. The default export receives ExtensionAPI ; it registers a command and an event handler. The handler narrows the tool type, checks its target, and returns Pi's supported blocking result. Pi loads TypeScript extensions directly, so this example needs no separate build step. This is a teaching example, not a filesystem sandbox : it does not cover shell commands, other tools, or alternate paths through symlinks. It also cannot tell whether an implementation respects the requirement. That semantic question is where the Jev experiment begins. In the real extension, a judgment starts with a specific event and a bounded evidence packet. The packet separates user requirements, observed tool results, actor claims, and hypotheses. A claim that “the tests passed” is different from the recorded output of a test command. The sequence is: Jev's output is input to a policy. It is not executable authority. A concern without a usable evidence anchor may produce no intervention. The source separates these responsibilities: src/policies/ https://github.com/krisitown/jev-loop-control/tree/v0.5.0-dev.1/src/policies holds policy decisions, while src/policy-live.ts https://github.com/krisitown/jev-loop-control/blob/v0.5.0-dev.1/src/policy-live.ts connects them to Pi. This makes it possible to test delivery with controlled judgments before asking whether hosted judgments are useful. The earlier supervisor inspected proposals broadly. The redesign asks narrower questions at distinct boundaries: | Policy | Trigger | Useful question | |---|---|---| | Context Governor | Actor offers a source-linked fact; memory is selected before a request. | What should we retain or bring back into view? | | Decision Gate | Actor reports a consequential decision. | Does this choice conflict with the requirement or lack critical evidence? | | Assumption Gate | Actor reports an assumption it is about to rely on. | Can we carry it forward, verify it cheaply, or do we need user input? | | Progress Governor | Actor reports progress, or enough tool results accumulate. | Are we learning something, or repeating an unproductive approach? | | Completion Gate | Pi reaches its completion boundary. | Does current evidence support the claimed result? | Decision and assumption checks depend on the actor reporting them through jev checkpoint . An unreported decision can still be missed. Progress and completion also have automatic triggers. These names describe this project's modules, not a standard taxonomy. For the parser example, tests that accept valid settings do not establish that unknown settings are rejected. A completion check needs evidence for the requested behavior, tied to the code revision being claimed complete. An old passing test cannot automatically clear a later edit. The video includes a planted conflict : the requirement says to reject unknown settings, while the proposed change accepts them. That demonstrates a control path. We also needed ordinary tasks, productive controls, and independent acceptance checks. The report https://zenodo.org/records/23046412 separates the studies because they answer different questions. They span development stages; this is not one benchmark of a single frozen version: | Study | Observed result | What it supports | |---|---|---| | 13 final-source Pi smoke scenarios | 10 of 12 Qwen cases passed their declared assertions. A separate scripted actor's dependent write was blocked. Two Qwen checks missed their original expectations and remain reported. | Controlled judgments can reach Pi and affect execution. This does not measure hosted-model usefulness. | | Three-task paired pilot | Both arms scored 14/15 acceptance checks. Jev logged 122 assessments and 76 interventions , with greater recorded time and token use. | More intervention did not produce a higher score in this pilot. | | Seeded policy calibration | At exploratory thresholds of 0.50 probability / 0.10 margin , only context selection met its gate: four repeated challenge observations and four controls. | A promising result on authored packets, not eight independent held-out tasks or proof of context savings. | | Short decision challenge and control | Plain Pi and a generic checklist completed both. Both Jev continuations timed out; three usable judgments produced no corrective effect. | Cleaner diagnostic prompts did not transfer reliably to these live continuations. | | Adaptive sample-project campaign | Mixed outcomes. The later B9 retry ended at 253/256 checks in both arms, with 30 Jev assessments and no applied interventions . | Useful traces to investigate; changing versions, different starting histories, and incomplete runs prevent a causal comparison. | In the three-task pilot, recorded input tokens were about 5.4× higher with Jev; actor time was about 1.6× higher. These are totals from a tiny exploratory study, not a forecast for other projects or proof that supervision caused every part of the overhead. Three lessons changed the design. Zero interventions is ambiguous: the actor may already have corrected itself, or the evaluator may have been unavailable or insufficiently grounded. Successful execution is not successful verification: a command can finish while its output reports failing tests. And a better prompt on clean examples is not enough: the question must still work with the noisy evidence a real run supplies. All five policies are implemented. General coding benefit remains unproven. Context selection is the most promising lead from these tests, but we have not demonstrated that it reduces the actual actor prompt while preserving held-out acceptance. The source, README, and configuration example https://github.com/krisitown/jev-loop-control/tree/v0.5.0-dev.1 are public. The version discussed here is 0.5.0-dev.1 , a development release, tested with Pi 0.87.1 and Node 22.19.0+ . Install the pinned version: pi install git:github.com/krisitown/jev-loop-control@v0.5.0-dev.1 export AI GATEWAY API KEY="your-vercel-ai-gateway-key" Use a Vercel AI Gateway key with access to typesafe-ai/jev . Start a new Pi process after exporting it. The extension does not automatically load .env files. Your actor can remain local; bounded assessment packets go to the configured Jev endpoint, by default Vercel AI Gateway's TypeSafe endpoint. For a first inspection, create jev-loop-control.config.json in your test project: { "mode": "observe", "budget": { "maxRequests": 20 }, "limits": { "maxAssessments": 20, "maxInterventionsPerTask": 4, "maxTerminalContinuations": 1 } } Then start pi and run: /jev-status /jev-doctor /jev-trace observe records judgments without applying guidance, blocks, or policy memory promotion. It still makes hosted calls. It also omits the enforcement-mode checkpoint instructions, so it is not a behaviorally identical rehearsal of every decision and assumption check. Without a configuration file, the package defaults to enforce , all five policies enabled, and unlimited requests, assessments, and interventions. The small limits above are deliberate choices for this walkthrough, not the package defaults. A limit being reached is not a successful verification. When you want to exercise supported effects, restart in enforcement mode: JEV LOOP CONTROL MODE=enforce pi The environment mode overrides the file; its budgets still apply. Give the agent a small task with a clear, independently testable contract. For example: Implement a settings parser that accepts port and logLevel , rejects every unknown key, and tests both behaviors. Before making a consequential design choice, call jev checkpoint separately and await the result. Cite available evidence sources and report verification only after the relevant tool has run. This is a normal task, not a guarantee of an intervention. A correctly proceeding actor is a useful control. Do not lower thresholds simply to manufacture a block. Use the trace directory reported by /jev-trace to inspect events.jsonl and, when enabled, the sanitized request/response artifacts. Follow event → evidence → judgment → effect → delivery → subsequent behavior . A recorded recommendation is different from a delivered block, and a delivered block is different from a better result. Project content in traces can still be sensitive even though credentials are scrubbed. I want to isolate context selection: start every arm from the same task and history, place an important fact early, and later require it. Add irrelevant and superseded facts so retaining everything cannot win by default. Compare plain history, simple deterministic selection or summaries, and Jev selection. Measure independent acceptance, the actual context sent to the actor , tokens, time, and harmful omissions. For decision checks, include a generic “check your approach” reminder as a baseline. If that reminder helps just as much, the reminder deserves the credit. The goal is an agent I need to interrupt less. That has to show up in outcomes, not just a busy supervisor log. 10.5281/zenodo.23046412 . It has not been peer reviewed. The full campaign traces are not all bundled with the extension repository. The report identifies the studies and their limitations; the repository provides the implementation and its tests. If you have tried supervising a coding agent, what did you measure: a missed bug caught, fewer repair rounds, less context, or mostly another source of latency? A small, reproducible comparison against a simpler alternative would be especially useful.