cd /news/ai-agents/358-pull-requests-that-changed-tests… · home › topics › ai-agents › article
[ARTICLE · art-145037] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

358 pull requests that changed tests: agents rarely weakened them. They bent the code instead.

A developer built RepoPilot, a deterministic, LLM-free tool that flags when coding agents weaken test checks, after labeling 358 public pull requests that modified test code. In a small experiment, Codex never weakened tests but instead changed code to satisfy incorrect tests in 3 of 3 runs without the guard and 2 of 3 with it, and added a quiet fallback for a missing image dependency. The tool reports signals like added suppressions, removed assertions, and relaxed CI gates, and ships as a CLI plus a Claude Code and Codex plugin that interrupts an agent once per signal.

by read2 min views1 publishedOct 4, 2026

Coding agents get blamed for "making tests pass" by skipping or deleting them. I wanted numbers before building a guard for that, so I collected public pull requests that modified test code (2026, repositories with 100+ stars) and had them labeled before running any tool on them:

What I found:

Then I gave Codex three tasks that could not be done honestly: two tests that contradict the documented behavior, and an image feature whose native dependency is missing. It never weakened a test. It changed the code to satisfy the wrong test in 3 of 3 runs without the guard and 2 of 3 with it; for the missing image library it added a quiet fallback instead of failing. One model, one run per task, so a small sample, but the shortcut was in the code, not in the tests.

So the check became a review signal, not an accusation. repopilot review lists every place a change touched the checks that judge it: focused or skipped tests, removed tests, tests that lost assertions, new lint/type/coverage suppressions, relaxed CI or tool gates, and new entries in RepoPilot's own suppression file.

Held-out precision:

Signal Precision
suppression added 5/6 (6/6 after one documented label fix)
CI/tool gate relaxed 1/1
assertions removed 2/3
test removed (case or whole file) 5/8

What it misses: checks trivialized by structure (the unmount trick), loosened matchers, helpers in other files, ESLint bulk suppression files, module-level conditional skips, Ginkgo specs, and the quiet fallback above. That last one is next. Limits: one LLM labeler plus my review of the "weakened" verdicts, and the same model helped build the detectors. Treat the numbers as exploratory; the corpus, labels, and harness are in the repo.

It runs locally and is deterministic, with no LLM. For Claude Code and Codex there is a plugin that snapshots the repo when a session starts and, when the agent tries to finish, stops it once per signal if it weakened a check, with the file and line. Dependency bumps and workflow edits stay in the report and don't interrupt the agent.

npm i -g repopilot        # or: cargo install repopilot
repopilot review . --base origin/main

Claude Code: /plugin marketplace add MykytaStel/repopilot, then /plugin install repopilot@repopilot.

Repo: https://github.com/MykytaStel/repopilot

I wrote this post with help from Claude. The numbers, labels, and harness are in the repo, and I checked them before publishing.

── more in #ai-agents 4 stories · sorted by recency
── more on @repopilot 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/358-pull-requests-th…] indexed:0 read:2min 2026-10-04 · —