{"slug": "make-no-mistakes-agent-usability-testing", "title": "Make no mistakes: agent usability testing", "summary": "SerpApi developed an agent usability testing (AUT) skill that uses a second AI agent to observe and score how well LLM agents follow agentic skills, achieving 100% task success with the skill versus 20% without on scholar citations, flight prices, and Maps reviews across 432 test runs. The approach, tested with Copilot CLI and SerpApi, helped self-fix documentation issues that manual review missed, including auth checks and response JSON paths.", "body_md": "What if we could observe how AI agents follow agentic skills? We can. Spawn another agent process in background, read its output, score against ground truth. TDD or User Testing for skill files.\n\nI was updating our [ serpapi-web-search skill](https://github.com/serpapi/skills/blob/master/skills/serpapi-web-search/SKILL.md) and wanted to know: does this specific documentation help LLM agents correctly solve the task?\n\n## Results\n\n### Fixes\n\nThis approach helped self-fixing several documentation issues:\n\n[Added auth check in](https://github.com/serpapi/skills/commit/7a8cfdf0b90fae52a925ce1c073d5f2cb852af45)`SKILL.md`\n\n- Fixed response JSON path\n[several](https://github.com/serpapi/skills/commit/43d8413)[times](https://github.com/serpapi/skills/commit/2780dba)\n\n### Test runs\n\nExample prompt: \"How many Google Maps reviews does Tartine Bakery have?\"\n\nGround truth: 5,915.\n\nWith the skill: agent calls SerpApi, gets 5,915.\n\nWithout: Opus refused, Fable launched Playwright (works 5/12 times), GPT guessed confidently wrong.\n\nFor the different task, GPT gets exact flight prices WITHOUT the skill ($299, matching API answer, 5/6 correct). Fable and Opus: 0/12.\n\nWe ran 432 of these.\n\n| Task | With | Without | Notes |\n|---|---|---|---|\n| Scholar citations | 100% | 28% | Without gives stale 265K vs exact 257K |\n| Flight prices | 100% | 14% | GPT gets it via web search; Fable/Opus can't |\n| Maps reviews | 100% | 19% | Fable uses Playwright; Opus refuses; GPT guesses |\n| Hotel prices | 100% | 0% | Refuses: \"I can't look up live hotel prices\" |\n| Image source | 100% | 0% | Hedges or guesses wrong domain |\n| Product price | 100% | 33% | Gives $449-499 estimates; actual $649 |\n| Shopping price | 67% | 0% | 1 auth failure; gives ~$198 vs $248 actual |\n| YouTube views | 100% | 92% | Zero lift. Web search returns exact counts |\n| App Store rating | 100% | 100% | Zero lift. Public API endpoint |\n| JWST mass | 97% | 97% | Control. Training data |\n| News headlines | 100% | 100% | Control. Any web search works |\n\nScholar + Flights + Maps: WITH 100%, WITHOUT 20%. Haiku provided correct answers only when skill was provided.\n\n## How it works\n\nThe [AUT (Agent Usability Testing) skill](https://github.com/serpapi/skills/blob/master/skills/agent-usability-test/SKILL.md) runs on two conditions:\n\n**WITH:** skill file loaded + tool available**WITHOUT**: no skill file, no tool\n\n`WITH`\n\nscores higher: the skill was useful. `WITHOUT`\n\nscores same: it wasn't useful.\n\nIt guides an agent through the steps:\n\n- Write tasks your tool should help with (and control tasks it shouldn't)\n- Run WITHOUT. Watch it fail. That's your red test.\n- Run WITH. Watch it pass. That's green.\n- Edit one line of docs, rerun. Did it break? That's refactor.\n- Repeat across models.\n\n## Lessons\n\nI did several mistakes during evaluation:\n\n- Compared real\n`$HOME`\n\n(all configs and extensions) vs stripped HOME. Not a fair comparison. - Running from the repo directory. Agent grepped the working dir and found skill files.\n- A second\n`serpapi`\n\nbinary in`~/.cargo/bin`\n\nsurvived`PATH`\n\nshadowing.\n\n## Agents didn't auto-discover tools\n\nI ran a third condition: tool on `PATH`\n\nand authenticated, but no skill file.\n\n0 out of 24 runs. Never typed `serpapi`\n\n. Never explored PATH. Never discovered it.\n\n## Conclusion\n\nWe tested this \"TDD for skills\" approach with [Copilot CLI](https://docs.github.com/en/copilot/github-copilot-in-the-cli) and [SerpApi](https://serpapi.com/search-api) and fixed several documentation issues, that the `/review`\n\non clean context didn't notice. Please give the it try and share your results.\n\n## Related:\n\nP.S. For transparency, below are all tasks during the tests:\n\n### Phase 1: Easy tasks (126 runs, all 100%/100%, zero lift)\n\n- \"What is the phone number for Tartine Bakery in San Francisco?\"\n- \"What is Apple's current stock price?\"\n- \"What is the price of the iPhone 15 Pro?\"\n- Other simple lookup questions that built-in web_search handles trivially\n\n### Phase 2: First hard tasks (90 runs, N=3)\n\n- \"Exact citation count for 'Attention Is All You Need' on Google Scholar\"\n- \"Cheapest nonstop flight JFK→LAX tomorrow, one-way\"\n- \"Price of DeWalt DCD771C2 drill on Google Shopping by retailer\"\n- \"Number of Google Maps reviews for Tartine Bakery, SF\"\n- \"Google Maps rating for Tartine Bakery, SF\"\n\n### Phase 3: Expanded 240-run retest (10 tasks, 6 engines)\n\n- Scholar: AlexNet citations, BERT citations\n- Flights: JFK→SFO one-way, SFO→NRT\n- Shopping: DeWalt drill price\n- Maps: Central Park review count, Tsukiji rating\n- Hotels: cheapest Kyoto hotel on specific date\n- App Store: Notion app rating\n\n### Phase 4: Final N=12 matrix (432 runs, 3 models)\n\n- \"How many citations does 'Attention Is All You Need' have on Google Scholar?\"\n- \"What is the cheapest nonstop flight from JFK to SFO on 2026-07-09, one-way?\"\n- \"How many views does the official Rick Astley Never Gonna Give You Up video have on YouTube?\"\n- \"How many Google Maps reviews does Tartine Bakery in San Francisco have?\"\n- \"What is the mass of the James Webb Space Telescope's primary mirror in kilograms?\"\n- \"What are the top 3 technology news headlines right now?\"\n\n### Phase 5: Expanded engines (36 runs, Opus only, N=3)\n\n- \"What is the price of the PlayStation 5 Disc Console Slim on Walmart?\"\n- \"What is the price of the cheapest Nintendo Switch OLED listing on eBay?\" (invalid - no stable ground truth)\n- \"What is the all-time rating of the Notion app on the Apple App Store?\"\n- \"What is the cheapest hotel in Kyoto for the night of July 15, 2026?\"\n- \"What website hosts the first Google Image result for 'Earthrise Apollo 8 photo'?\"\n- \"What is the current price of Sony WH-1000XM5 headphones on Google Shopping?\"", "url": "https://wpnews.pro/news/make-no-mistakes-agent-usability-testing", "canonical_source": "https://serpapi.com/blog/make-no-mistakes-agent-usability-testing/", "published_at": "2026-07-06 11:22:52+00:00", "updated_at": "2026-07-21 18:20:09.537666+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools"], "entities": ["SerpApi", "Copilot CLI", "GitHub Copilot", "Playwright", "Opus", "Fable", "GPT", "Haiku"], "alternates": {"html": "https://wpnews.pro/news/make-no-mistakes-agent-usability-testing", "markdown": "https://wpnews.pro/news/make-no-mistakes-agent-usability-testing.md", "text": "https://wpnews.pro/news/make-no-mistakes-agent-usability-testing.txt", "jsonld": "https://wpnews.pro/news/make-no-mistakes-agent-usability-testing.jsonld"}}