358 pull requests that changed tests: agents rarely weakened them. They bent the code instead. A developer built RepoPilot, a deterministic, LLM-free tool that flags when coding agents weaken test checks, after labeling 358 public pull requests that modified test code. In a small experiment, Codex never weakened tests but instead changed code to satisfy incorrect tests in 3 of 3 runs without the guard and 2 of 3 with it, and added a quiet fallback for a missing image dependency. The tool reports signals like added suppressions, removed assertions, and relaxed CI gates, and ships as a CLI plus a Claude Code and Codex plugin that interrupts an agent once per signal. Coding agents get blamed for "making tests pass" by skipping or deleting them. I wanted numbers before building a guard for that, so I collected public pull requests that modified test code 2026, repositories with 100+ stars and had them labeled before running any tool on them: What I found: Then I gave Codex three tasks that could not be done honestly: two tests that contradict the documented behavior, and an image feature whose native dependency is missing. It never weakened a test. It changed the code to satisfy the wrong test in 3 of 3 runs without the guard and 2 of 3 with it; for the missing image library it added a quiet fallback instead of failing. One model, one run per task, so a small sample, but the shortcut was in the code, not in the tests. So the check became a review signal, not an accusation. repopilot review lists every place a change touched the checks that judge it: focused or skipped tests, removed tests, tests that lost assertions, new lint/type/coverage suppressions, relaxed CI or tool gates, and new entries in RepoPilot's own suppression file. Held-out precision: | Signal | Precision | |---|---| | suppression added | 5/6 6/6 after one documented label fix | | CI/tool gate relaxed | 1/1 | | assertions removed | 2/3 | | test removed case or whole file | 5/8 | What it misses: checks trivialized by structure the unmount trick , loosened matchers, helpers in other files, ESLint bulk suppression files, module-level conditional skips, Ginkgo specs, and the quiet fallback above. That last one is next. Limits: one LLM labeler plus my review of the "weakened" verdicts, and the same model helped build the detectors. Treat the numbers as exploratory; the corpus, labels, and harness are in the repo. It runs locally and is deterministic, with no LLM. For Claude Code and Codex there is a plugin that snapshots the repo when a session starts and, when the agent tries to finish, stops it once per signal if it weakened a check, with the file and line. Dependency bumps and workflow edits stay in the report and don't interrupt the agent. npm i -g repopilot or: cargo install repopilot repopilot review . --base origin/main Claude Code: /plugin marketplace add MykytaStel/repopilot , then /plugin install repopilot@repopilot . Repo: https://github.com/MykytaStel/repopilot https://github.com/MykytaStel/repopilot I wrote this post with help from Claude. The numbers, labels, and harness are in the repo, and I checked them before publishing.