Handing an agent's browser clicks to Jev A developer built jev-browser-wingman, a middleware layer that sits between an AI agent and its browser tool (such as Playwright MCP) and uses TypeSafe's Jev to grade how well each page element fits a step, acting directly when the match is clear and handing the step back to the agent when it is not. The project reached version 0.3.0 after roughly twenty validation rounds in a week, and the team found that hiding the browser server's own click, type, select, key, hover, upload, navigate, back, scroll and script tools ("hand-off mode") was the only reliable way to get agents to use it, after prompt rewrites failed. A safety bug that could submit a form twice when a submit left the page at the same URL was fixed, and wingman was about 35% cheaper on a seven-page benchmark task once cache-warmth differences were removed. Updated 2026-10-08 with 0.3.3 benchmark results. Every click an AI agent makes in a browser costs a full model turn, and so does every keystroke, scroll and wait. We wanted to find out how much of that page work TypeSafe's Jev could take over. This post tells how jev-browser-wingman wingman for short got to version 0.3.0: the ideas that failed, a bug that submitted a form twice, and a clever trick that the numbers talked us out of. A validation round, for us, was a batch of fixes followed by a benchmark run, and we did about twenty of them in a week. Results from the current version are at the end. We built it with AI coding agents. One agent planned each round and coordinated the work. Others wrote the code against a fixed spec, and separate agents were told to break each change rather than approve it. The idea: a loop in front of the browser tool Wingman sits between the agent and its browser tool, usually an MCP server such as Playwright MCP. MCP, the Model Context Protocol, is how agents talk to tool servers, and Playwright is a browser-automation library. For each step, wingman reads the page and asks Jev to grade how well each element fits. When one element clearly fits, wingman acts, checks what happened and moves on. When Jev is unsure, wingman hands the step back to the agent. Asking the agent to delegate did not work The first problem was behavioral, not technical. We assumed that a good tool description and a clear prompt would make the agent use wingman. On our longest task, a seven-page chain described in the next section, they did not. The agent either ignored the tool or never chose it while its regular browser tools were still visible. We rewrote the tool description, then added a line telling the agent to engage. Neither fixed it. What worked was removing the choice. In what we now call hand-off mode, wingman hides the browser server's own click, type, select, key, hover, upload, navigate, back, scroll and script tools. If you want an agent to use a tool, do not rely on persuasion in the prompt. Change what it can see. One hard task: seven pages on a live practice site We needed one task that could hurt us. We chose a chain of seven pages on the-internet, a public site built for practicing browser automation. The pages cover checkboxes, a dropdown, adding and removing elements, inputs, a forgot-password form, dynamic loading and status codes. We measured three runs at each stage, starting from a low point that was worse than the round before it. The fixes covered stuck-loop recovery, deciding a step had finished once a click left the page, and ending a call after an action had run. That was not a steady climb. In the final benchmark, on local copies of the pages, we did not close the time gap, although wingman was about 35% cheaper on this task with the cache-warmth difference removed explained under Results . A safety bug: the form that submitted twice The scariest defect came out of this work. After a form submit that left the page at the same address, wingman could not tell that the step had finished. The practice site's forgot-password form does exactly this: it answers with an error page at the same URL. This happened in several early runs. On a real checkout it would be a duplicate order. We fixed it. That is a fix, not a guarantee, and the README still tells callers to look at the page before retrying a handed-back step. Real pages exposed gaps, and one obvious idea failed Passing our unit tests had given us false comfort. So we ran two rounds on five unfamiliar task types. These found six defects in the loop that the unit tests had missed. That result is why a standing benchmark of 17 tasks now exists. Each fix lands with a test we have watched fail without it. A test that has never failed proves nothing. One calibration finding shaped the design too. Jev graded the element choice honestly, but our question set had no way to ask whether the element was covered. Then came the idea we cut. Some candidate elements never showed up because they had no click handler we could see, and a tempting fix was to treat anything styled with a pointer cursor as clickable. It looked obvious. We measured it on three live news sites. The tool still cannot click an element that reacts only through a listener it cannot see. That limit is documented. A known gap that we can explain beats a heuristic that makes pages worse. When the test site became the problem The public practice site started to stall, which made it hard to tell a wingman regression from a site hiccup. Results today, and the limits These results come from a run on the 0.3.2 code. 0.3.3 changes only the confirm-gate and wait-for-disappearance paths, which this run did not exercise the gate was off and the wait shortcut never fired , so we have not re-run it. The benchmark has 17 tasks, each run twice with wingman and twice with Playwright alone, 68 runs in total. Wingman completed 34 of 34 runs and Playwright alone 32 of 34. Both of Playwright alone's misses were on the file-upload task. We have not diagnosed why, and with one task and two runs we do not read it as a reliability result. Total cost over all runs was $3.75 with wingman and $5.35 with Playwright alone, 30% lower. We do not quote the median run cost, because costs split by whether the agent's prompt cache was already warm. Playwright alone started cold in 17 of its 34 runs and wingman in 13. Pricing every cache write at the cache-read rate largely removes that effect: the totals become $2.22 and $3.22, 31% lower. Wingman was cheaper on 16 of 17 tasks and dearer only on the confirm dialog, by 2% before and 3% after the cache adjustment. The median run took 16.0 seconds with wingman and 15.5 with Playwright alone, so wingman was 0.5 seconds slower. It was faster on 8 tasks and slower on 9, by 4 seconds or more on 3: the seven-page chain, the demo-shop checkout and the dropdown. The mean run was 1.2 seconds slower 18.0 s against 16.9 . Wingman's own work reading the page, asking Jev, acting, waiting for the page to settle was a small share of most runs. It took under 10% of the run time in both runs on 9 of 17 tasks, and 39% to 44% on the seven-page task. On the wait-for-hidden-text task it was 34% to 36%, and that task is mostly waiting. Waiting for hidden text was our worst result at 0.3.0 41.9 seconds against 20.5 . In this run wingman was faster in both runs, averaging 16.8 seconds against 22.1; Playwright alone's two runs took 16.9 and 27.3. Wingman handed a step back to the agent in 1 of 34 runs, and that run still completed. The path was 17 of 34 runs at 0.3.0, 4 of 34 at 0.3.1 and 1 of 34 now, so most of the drop came in 0.3.1. Counting any return that was not a finished task, such as a login page or a page error, the figure for this run is 4 of 34. The release notes list what changed since 0.3.0: github.com/coderexpert123/jev-browser-wingman/releases https://github.com/coderexpert123/jev-browser-wingman/releases Some caveats apply. There are two runs per task, so the numbers are indicative. The wingman setup's prompt has three extra sentences about using browse step, and they explain how to call it. Our earlier finding that a prompt alone did not win delegation came from runs where the agent's own click tools were still visible. The 0.3.0 run's prompt had two of these sentences; the third, telling the agent to skip a step that opens the page already open, was added later and removed the 404-page task's hand-backs. Playwright alone gets no tool-specific guidance. The turn cap is 30 for wingman and 40 for Playwright alone, though no run made more than 20 tool calls. Times rose on both setups compared with the 0.3.0 and 0.3.1 runs, so compare times within one run only. The calling agent was a fast model served through a proxy and reached through Claude Code's sonnet alias, so this is not a direct measurement of Anthropic's Sonnet model. Its tokens are priced at Sonnet list rates, so treat the dollars as an index: the ratio between the two setups is the result. The per-task table is in the README under "Results by task": github.com/coderexpert123/jev-browser-wingman results-by-task https://github.com/coderexpert123/jev-browser-wingman results-by-task The limits are real. Wingman cannot see inside iframes or shadow-DOM content. Right-clicks, modifier-clicks and unusual keys need the agent's own script tool, which you must choose to keep. What we would tell other builders To try wingman, the code is at github.com/coderexpert123/jev-browser-wingman https://github.com/coderexpert123/jev-browser-wingman . Install it with npm install -g jev-browser-wingman . The README has recipes for Claude Code with Playwright MCP and for other browsing servers.