# 358 pull requests that changed tests: agents rarely weakened them. They bent the code instead.

> Source: <https://dev.to/cherven/358-pull-requests-that-changed-tests-agents-rarely-weakened-them-they-bent-the-code-instead-3ld>
> Published: 2026-10-04 22:02:51+00:00

Coding agents get blamed for "making tests pass" by skipping or deleting them. I wanted numbers before building a guard for that, so I collected public pull requests that modified test code (2026, repositories with 100+ stars) and had them labeled before running any tool on them:

What I found:

Then I gave Codex three tasks that could not be done honestly: two tests that contradict the documented behavior, and an image feature whose native dependency is missing. It never weakened a test. It changed the code to satisfy the wrong test in 3 of 3 runs without the guard and 2 of 3 with it; for the missing image library it added a quiet fallback instead of failing. One model, one run per task, so a small sample, but the shortcut was in the code, not in the tests.

So the check became a review signal, not an accusation. `repopilot review` lists every place a change touched the checks that judge it: focused or skipped tests, removed tests, tests that lost assertions, new lint/type/coverage suppressions, relaxed CI or tool gates, and new entries in RepoPilot's own suppression file.

Held-out precision:

| Signal | Precision | 
|---|---|
| suppression added | 5/6 (6/6 after one documented label fix) | 
| CI/tool gate relaxed | 1/1 | 
| assertions removed | 2/3 | 
| test removed (case or whole file) | 5/8 | 

What it misses: checks trivialized by structure (the unmount trick), loosened matchers, helpers in other files, ESLint bulk suppression files, module-level conditional skips, Ginkgo specs, and the quiet fallback above. That last one is next. Limits: one LLM labeler plus my review of the "weakened" verdicts, and the same model helped build the detectors. Treat the numbers as exploratory; the corpus, labels, and harness are in the repo.

It runs locally and is deterministic, with no LLM. For Claude Code and Codex there is a plugin that snapshots the repo when a session starts and, when the agent tries to finish, stops it once per signal if it weakened a check, with the file and line. Dependency bumps and workflow edits stay in the report and don't interrupt the agent.

```
npm i -g repopilot        # or: cargo install repopilot
repopilot review . --base origin/main
```

Claude Code: `/plugin marketplace add MykytaStel/repopilot`, then `/plugin install repopilot@repopilot`.

Repo: [https://github.com/MykytaStel/repopilot](https://github.com/MykytaStel/repopilot)

I wrote this post with help from Claude. The numbers, labels, and harness are in the repo, and I checked them before publishing.
