tested OpenRig as a way to give a coding task to one agent and return a concrete review from another. The useful outcome was a working local TODO page, an executed review, and task records that connected the two. Human intervention remained part of the workflow.
Environment and scope
This experiment ran on October 2, 2026, on one Mac running macOS 27.0, with Node.js 22.23.2 and tmux 3.7c. I pinned @openrig/cli to 0.6.4. Two native Codex sessions used gpt-6-astra, with owner and checker roles. I did not test heterogeneous models or a Claude provider.
The current upstream README may describe behavior beyond the published version I tested. Always compare those instructions with the installed CLI version. The official repository is https://github.com/mvschwarz/openrig.
What the ten installation rounds establish
Each final round used a fresh npm installation prefix and retained its own log and genuine screen recording. System dependencies were prepared once on the shared machine. All ten installations succeeded. Each round also passed version inspection, a preview of the first-project rig specification, and setup --dry-run, for thirty successful checks in total.
The first round used an empty cache and took 8.513 seconds for the installation process. The subsequent nine rounds used a warm cache and took 1.724 to 2.303 seconds, with a median of 1.837 seconds. These are local installation-process measurements. They exclude account setup, model execution, browser work and operator troubleshooting. The warm rounds cannot serve as an uncached performance benchmark.
Fresh prefixes on one host do not establish cross-platform compatibility. Also, setup --dry-run is a preview and preflight check. The experiment did not apply full workstation setup ten times.
For a reproducible fresh-prefix check, I would retain this shape and keep the system dependencies constant:
npm install --prefix "$PWD/prefix" @openrig/[cli@0.6.4](mailto:cli@0.6.4)
"$PWD/prefix/node_modules/.bin/rig" --version
"$PWD/prefix/node_modules/.bin/rig" specs preview first-project --kind rig
"$PWD/prefix/node_modules/.bin/rig" setup --dry-run This is the inspection phase. It does not launch agents, grant permissions or prove that provider authentication works. Keep subsequent agent launch and workspace changes within the project you intend to test, using the upstream instructions for your selected provider.
The task and its acceptance contract
The owner produced a local static TODO page and a browser test harness. The required behavior was deliberately small: add tasks, toggle completion in both directions, and preserve state after a reload. Empty and whitespace-only input had to be rejected. HTML-like input had to remain text.
I exercised the page in real Chrome. Adding a Chinese task worked. Completion survived a reload, and changing it back also persisted. An HTML payload remained literal text: no image node was created and no dialog appeared. Long text and emoji retained their content.
I then checked storage failure paths. Corrupt JSON produced a warning while the page remained usable. A malformed stored schema was rejected. Simulated read and write failures generated the appropriate warnings, and operations that could still proceed remained available. Ten extended operator checks passed. Their coverage is specific to this page and this browser session; it is not a comprehensive security assessment.
What the second agent actually reviewed
The checker read the source and verified the candidate files. It independently reran the developer's existing harness in real Chrome and returned a QA report. The bounded contract passed, and no blocking defect was identified. There was no invented bug-fix cycle in this result.
Independent execution and independent test design are different claims. The checker reused the existing harness. I did not establish that it built a separate adversarial suite. The operator's extended tests were a distinct verification step, and should not be attributed to the checker.
The queue also contained a claimed development task, a recorded review handoff and a returned report. These records made it possible to connect ownership and review to the produced artifacts rather than infer completion from chat messages alone.
Operator intervention that belongs in the report
An initial native-model configuration pointed to an inactive local endpoint. I corrected my environment before the sessions could proceed. A separate isolated-prefix issue required explicit command paths and instance-directory information. Those were operator/environment problems, and this run does not justify labeling them OpenRig defects.
Browser launch and access to the local service still needed explicit approvals. The isolated instance used a no-kernel mode, and Codex activity hooks were disabled for this run. Therefore, neither kernel behavior nor those hooks are validated by the result. A single existing-session resume was observed; long-term persistence remains untested.
What I would test next
The next useful experiment would give the reviewer its acceptance contract before it reads the developer's report, then ask it to extend the suite. I would also test browser restarts, another browser, longer-lived stored data and failure recovery. Larger projects, more agents and mixed providers need their own runs.
For this small task, OpenRig provided usable task ownership and a review handoff. The installation evidence, browser behavior and review record support that limited conclusion. I cannot report a productivity percentage because this experiment did not include a timed control workflow. Sources and artifacts
Upstream project: https://github.com/mvschwarz/openrig The companion evidence for this episode consists of ten final installation logs and recordings, the static page, the existing browser harness, the QA report and the operator's extended-check results. The release handoff retains these artifacts locally. No production deployment or public test service was created.
Authored for the Cao Ge Tests AI series.