{"slug": "i-built-agent-review-studio-a-local-first-workbench-for-agent-harness", "title": "I built Agent Review Studio: a local-first workbench for agent harness evaluations", "summary": "An engineer at Chase built Agent Review Studio, an open-source, local-first workbench for evaluating and refining AI agents and agent harnesses. The tool lets engineers inspect execution traces, classify evidence, and turn reviewed failures into golden evaluation cases or regression tests. Version 1.5.0 is available under Apache-2.0.", "body_md": "I have been building Chaser Agent and other AI systems across the Chase ecosystem. As those systems became more capable, I needed a better way to inspect what an agent actually did—not just look at its final answer.\n\nSo I built **Agent Review Studio**.\n\nIt is an open-source, local-first system for evaluating and refining AI agents and agent harnesses.\n\nAn agent run can produce a final answer, source files, extracted claims, proposed actions, memory candidates and a complete execution trace. Looking only at the final response hides most of the engineering evidence.\n\nI wanted one review flow that could answer practical questions:\n\nAgent Review Studio lets an engineer:\n\nThe important part is what happens next. A reviewed failure can become a golden evaluation case or regression test. We can update a prompt, retrieval system, tool policy, memory rule or workflow, run the same task again and compare the result with the original baseline.\n\nThe launch demo uses a real Chaser Agent research run built around a public Cloudflare source. Agent Review Studio opens the run's files, places extracted claims beside the retained evidence and lets the operator classify what was correct, weak, missing or irrelevant.\n\nThe example labels shown in the video are an unsaved demonstration. The +17 result is an automated comparison, and human review is still pending. I kept those boundaries visible because an automated score should not be mistaken for human approval.\n\nChaser Agent is only one workspace. Fresh installations of the Studio contain no preloaded agent. Other engineers can name the system they are building, describe its objective and import their own datasets and runs.\n\nThis is agent evaluation, human labelling, evidence curation and harness refinement.\n\nThe Studio does not automatically change model weights. It creates the evaluation infrastructure and trusted improvement data needed to refine prompts, tools, retrieval, memory and orchestration. Reviewed examples can later become candidates for a separate, governed model fine-tuning pipeline.\n\nVersion 1.5.0 is open source under Apache-2.0.\n\nI would value feedback from other people building agents: **what evidence do you require before deciding that a new harness version is genuinely better?**\n\n*Disclosure: I used AI assistance while preparing and editing this launch article. I reviewed the claims, product boundaries and final publication myself.*", "url": "https://wpnews.pro/news/i-built-agent-review-studio-a-local-first-workbench-for-agent-harness", "canonical_source": "https://dev.to/chaseintech/i-built-agent-review-studio-a-local-first-workbench-for-agent-harness-evaluations-158p", "published_at": "2026-09-03 11:08:36+00:00", "updated_at": "2026-09-03 11:24:55.153915+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-tools", "mlops"], "entities": ["Agent Review Studio", "Chaser Agent", "Chase", "Cloudflare"], "alternates": {"html": "https://wpnews.pro/news/i-built-agent-review-studio-a-local-first-workbench-for-agent-harness", "markdown": "https://wpnews.pro/news/i-built-agent-review-studio-a-local-first-workbench-for-agent-harness.md", "text": "https://wpnews.pro/news/i-built-agent-review-studio-a-local-first-workbench-for-agent-harness.txt", "jsonld": "https://wpnews.pro/news/i-built-agent-review-studio-a-local-first-workbench-for-agent-harness.jsonld"}}