{"slug": "my-code-reviewer-scored-a-nonexistent-directory-100-100-and-exited-0", "title": "My Code Reviewer Scored a Nonexistent Directory 100/100 and Exited 0", "summary": "A developer audited 34 agent skill packages and found that a code-review tool returned a 100/100 health score and exit code 0 after being passed a nonexistent path, scanning zero files while reporting 'safe to merge.' The audit also produced false negatives on five packages whose self-tests use a --selftest flag rather than a positional subcommand, revealing that two competing self-test conventions across the folder make harness invocation unreliable. The developer argues that a tool returning the same success status for 'nothing wrong' and 'nothing examined' cannot serve as a CI gate.", "body_md": "Every skill package I ship — for Claude Code, Cursor, Codex, and a couple of agent marketplaces — has to clear one gate before publishing: its own `selftest` must exit 0.\n\nSo this week I ran an audit over the 34 packages sitting in my skills folder to find out how many actually have that gate. The audit broke twice, in two different ways. Both breakages are the same bug wearing different clothes, and it's the bug that makes agent tooling untrustworthy the moment you put it in CI:\n\n**a success signal that isn't attached to work actually done.**\n\nHere's the whole thing, including the raw output.\n\nThe script was not clever. Walk each package, find a script mentioning `selftest`, run `python scripts/<script>.py selftest`, read the exit code.\n\nResult:\n\nThe five failures looked like this (output is Chinese, I'll translate):\n\n```\nFAKE(2) license-compliance-checker :: 错误: not a directory: ...\\license-compliance-checker\\selftest\nFAKE(2) meeting-minutes-assistant   :: 找不到文件：selftest\nerror: not a directory: ...\\selftest\nfile not found: selftest\n```\n\nThose five tools were fine. **My harness was wrong.** They implement the self-test as a `--selftest` flag, not as a positional subcommand. Five working self-tests (30/30, 54/54, 16/16, 40/40, 77/77 assertions) were reported as failures because the invocation contract was never part of the contract.\n\nThat's a false negative. Annoying, but it fails safe — the gate stays closed.\n\nOne of those five is a code-review tool. Called with a positional argument, it does not complain. It treats the argument as a path:\n\n``` bash\n$ python scripts/code_review.py selftest\n跳过不存在的路径: selftest\n# 代码审查报告\n- 审查范围: 0 个文件，0 行\n- 发现问题: 0 个（致命 0 / 高 0 / 中 0 / 低 0）\n- 健康分: 100/100 — ✅ 可以合并（按低优先级项择机清理）\nEXIT=0\nskipped nonexistent path: selftest\n# Code review report\n- scope: 0 files, 0 lines\n- findings: 0 (critical 0 / high 0 / medium 0 / low 0)\n- health score: 100/100 — ✅ safe to merge\nEXIT=0\n```\n\nRead that again. The tool skipped a path that doesn't exist, scanned zero files, handed out a perfect score and the string \"safe to merge\", and returned success.\n\nIts actual `--selftest` works fine — 54 assertions, 38 rules, 23 of them firing on dirty samples. The bug only shows up in scan mode, and only when the input is empty, missing or malformed. Which is exactly the state a path variable is in when a shell glob expands to nothing, a CI job runs on a docs-only diff, or an agent passes the wrong parameter name.\n\nNow imagine the two ways this gets wired up:\n\n`code_review.py --selftest && deploy` — fine.`code_review.py $CHANGED_FILES` — with `$CHANGED_FILES` empty, you get \"100/100, safe to merge\" on a diff nobody looked at, green pipeline, exit 0.\nAnd if there's an LLM agent on top reading that output, it will faithfully report \"review passed.\" Not because the model is sloppy — because the tool lied first.\n\n**\"Nothing wrong\" and \"nothing examined\" are not the same result, and a tool that returns the same status for both cannot be part of a gate.**\n\nAcross 34 packages:\n\n| self-test contract | packages | \n|---|---|\n| `--selftest` flag | 5 | \n| `selftest` positional subcommand | 12 | \n| no self-test | 17 | \n\nTwo conventions in one folder. Any harness that guesses will produce false negatives (safe) and — the dangerous half — false positives, since a wrong invocation is silently treated as *data* rather than as an error. That's not a hypothetical: it's the output above.\n\nAssertion counts in the 17 that do have a self-test: 129, 85, 77, 71, 54, 44, 44, 40, 30, 26, 21, 19, 18, 16, 15, 12, plus one boolean PASS. About 700 assertions total, and the useful ones are all two-sided:\n\nWithout the second half, a linter that flags everything looks perfect. Without the first half, a linter that flags nothing looks perfect. Exit-code-only gates cannot tell those apart.\n\nConcrete, from the 129-assertion SQL inspector:\n\n`DROP TABLE` is a finding; `DROP TABLE IF EXISTS` is not.`'no where clause here'` inside a string literal is not a missing WHERE clause.`SELECT *` inside a comment is not a finding.`os.environ[\"DB_PASSWORD\"]` is not a hardcoded credential; a literal password is.`BEGIN ... END` in a PL/pgSQL function body is not an unclosed transaction.\nEvery one of those negative cases exists because the naive version of the rule got it wrong first.\n\n`SKILL.md`, and have the harness read it from metadata instead of pattern-matching the source. Guessing is how you get the false negative; silent acceptance is how you get the false positive.`evaluation_cases.md` next to the source. A gate that trusts a previous run's report is not a gate.\nAnd one more, because it's the failure mode I keep hitting: **state what the tool does not do.** The review tool prints \"static rules can only disprove, not prove — still verify permissions, concurrency and money precision by hand.\" That sentence is the most honest thing in the package. A gate that doesn't declare its blind spots gets trusted for them.\n\nAdding self-tests to them is mechanical now — the templates exist, and the ones already done took about an hour of real assertions each, not real research. I'm starting with the packages that get wired into automation, because a missing self-test in a package nobody runs is a backlog item, while a zero-input success path in a pre-commit hook is an incident waiting for a quiet Friday.\n\nIf you maintain skills, MCP servers, or CLI wrappers that agents call: the highest-value 30 minutes you can spend today is running your tool with an empty input and reading what it says. If it says \"safe,\" \"passed,\" or \"0 issues,\" you've found the bug.\n\n**Where this stuff lives**\n\nIf you've hit a zero-input success path in a tool you rely on, I'd like the exact output — that's the most useful kind of bug report.", "url": "https://wpnews.pro/news/my-code-reviewer-scored-a-nonexistent-directory-100-100-and-exited-0", "canonical_source": "https://dev.to/felixwang007/my-code-reviewer-scored-a-nonexistent-directory-100100-and-exited-0-41m0", "published_at": "2026-10-01 04:01:42+00:00", "updated_at": "2026-10-01 04:16:38.790552+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "mlops"], "entities": ["Claude Code", "Cursor", "Codex"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/my-code-reviewer-scored-a-nonexistent-directory-100-100-and-exited-0", "markdown": "https://wpnews.pro/news/my-code-reviewer-scored-a-nonexistent-directory-100-100-and-exited-0.md", "text": "https://wpnews.pro/news/my-code-reviewer-scored-a-nonexistent-directory-100-100-and-exited-0.txt", "jsonld": "https://wpnews.pro/news/my-code-reviewer-scored-a-nonexistent-directory-100-100-and-exited-0.jsonld"}}