{"slug": "fakegreen-catch-ai-coding-agents-faking-a-green-build-no-llm", "title": "Fakegreen: Catch AI coding agents faking a green build (no LLM)", "summary": "A new open-source tool called fakegreen scans a git diff to detect when AI coding agents fake a passing build, flagging moves such as skipping tests, weakening assertions, deleting test files, adding @ts-ignore, branching on NODE_ENV === 'test', or appending `|| true` to CI commands. The tool runs via `npx fakegreen`, requires Node.js 18+ and git, uses no LLM or API key, has zero runtime dependencies, and typically finishes in under 100 ms; in a demo run against a sample repo it reported 5 high and 1 medium findings in 41 ms and exited with code 1. It is aimed at teams using Claude Code, Codex, Cursor, Gemini CLI, or Aider who want a deterministic pre-commit or end-of-turn tripwire that blocks an agent before it can claim \"All tests pass ✅\".", "body_md": "**One command scans your agent's diff and flags every way it faked a green build.**\n\n- Teams shipping with Claude Code, Codex, Cursor, Gemini CLI, or Aider who want a hard stop on fake-green diffs\n- Maintainers who want a deterministic CI / pre-commit / end-of-turn tripwire — no LLM, no API key, zero runtime deps\n- Anyone tired of agents that `.skip` tests, weaken assertions, or append`|| true` and then claim*\"All tests pass ✅\"*\n\n⭐ **If fakegreen catches a fake-green commit for you, [star the repo](https://github.com/fitzyracing1/fakegreen)** — it helps other teams find the tripwire.\n\n```\nnpx fakegreen\n```\n\nNo install. Needs Node.js 18+ and `git` on your `PATH`. Pin with `npm i -D fakegreen` when you're ready.\n\n|  | LLM code review | fakegreen | \n|---|---|---|\n| **Speed & cost** | Seconds–minutes, API spend | Usually **under 100 ms** , free, offline | \n| **Determinism** | Can vary; easy to talk around | Same answer every time | \n| **Setup** | API key + prompts | Zero runtime deps — `npx fakegreen` | \n\nCoding agents (Claude Code, Codex, Cursor, Gemini CLI, Aider) are rewarded for \"tests pass\". Sometimes they get there by\ndeleting the test, slapping `.skip` on it, swapping `toBe(42)` for `toBeDefined()`, adding `@ts-ignore`, teaching the code\nto detect `NODE_ENV === 'test'`, or appending `|| true` to the CI step. Then they say *\"All tests pass ✅\"*.\n\n`fakegreen` reads the git diff and catches those moves. It is **deterministic**, needs **no LLM and no API key**, has\n**zero runtime dependencies**, and usually finishes in **under 100 ms**. Hook it into your agent's end-of-turn event so the\nagent gets blocked and told what it did *before* it can claim it's done.\n\nThe recording above ([`docs/demo.cast`](https://github.com/fitzyracing1/fakegreen/blob/main/docs/demo.cast), [` docs/demo.gif`](https://github.com/fitzyracing1/fakegreen/blob/main/docs/demo.gif), [` docs/demo.png`](https://github.com/fitzyracing1/fakegreen/blob/main/docs/demo.png))\nis a real run against a sample repo built by [`scripts/make-demo-repo.sh`](https://github.com/fitzyracing1/fakegreen/blob/main/scripts/make-demo-repo.sh). An honest commit adds a\ncart module with tests. Then an \"agent\" commit titled `fix: make the test suite pass` skips one test, weakens an\nassertion, deletes a test file, adds `@ts-ignore` plus a `NODE_ENV === 'test'` shortcut, and appends `|| true` to CI:\n\n``` bash\n$ fakegreen --last-commit\nfakegreen · last commit (8797f43) · 4/4 files · +5 −13\n\n   HIGH  .github/workflows/ci.yml:10  ci-failure-ignored\n         Failure ignored with `|| true` on a test/check command\n         │ + - run: npm test || true\n\n   MED   src/cart.ts:4  suppression-added\n         Checker silenced with @ts-ignore\n         │ + // @ts-ignore\n\n   HIGH  src/cart.ts:5  test-env-special-case\n         Non-test code branches on the test env (NODE_ENV === 'test'); tests skip the real path\n         │ + if (process.env.NODE_ENV === 'test') return items.length ? 8 : 0;\n\n   HIGH  test/cart.test.ts:9  test-skipped\n         Test skipped with .skip\n         │ + it.skip('applies SAVE10', () => {\n\n   HIGH  test/cart.test.ts:14  assertion-weakened\n         Specific assertion replaced with a vague one\n         │ - expect(total([])).toBe(0);\n         │ + expect(total([])).toBeDefined();\n\n   HIGH  test/checkout.test.ts  test-file-deleted\n         Test file deleted (2 test cases removed)\n         │ - it('rejects negative quantities', () => {\n\n  5 high · 1 medium · 0 low  ✖ fake green detected (fail-on: high) · 41ms\n```\n\nExit code `1`. Replay it with `asciinema play docs/demo.cast`, or rebuild it with\n`scripts/make-demo-repo.sh /tmp/fakegreen-demo && cd /tmp/fakegreen-demo && fakegreen --last-commit`.\n\n```\nnpx fakegreen                 # run once, no install\nnpm i -D fakegreen            # or pin it in a project\n```\n\nYou need Node.js 18 or newer and `git` on your `PATH`.\n\n```\nfakegreen                     # uncommitted + staged changes vs HEAD, plus untracked files (default)\nfakegreen --staged            # only what is staged (good for pre-commit)\nfakegreen --base origin/main  # everything since the merge-base with main, including uncommitted work (good for PRs)\nfakegreen --last-commit       # just HEAD~1..HEAD\nfakegreen --commit <sha>      # one specific commit\ngit diff main... | fakegreen --diff -   # any unified diff from stdin or a file\n\nfakegreen --json              # machine-readable findings\nfakegreen --sarif > fg.sarif  # SARIF 2.1.0 for GitHub code scanning and IDEs\nfakegreen --format github     # ::error annotations for GitHub Actions\nfakegreen --fail-on medium    # exit 1 on medium or higher (default: high; `none` never fails)\nfakegreen --min-severity medium  # hide low-severity findings\nfakegreen rules               # list every rule\n```\n\nExit codes: `0` clean (or below `--fail-on`), `1` findings at or above `--fail-on`, `2` usage or git error.\n\nEach finding includes a severity, `file:line`, a rule id, a short explanation, and the offending `+`/`-` lines.\n`--json` output looks like this:\n\n```\n{\n  \"tool\": \"fakegreen\", \"version\": \"0.1.0\", \"source\": \"last commit (8797f43)\", \"failed\": true,\n  \"summary\": { \"high\": 5, \"medium\": 1, \"low\": 0, \"total\": 6, \"files\": 4, \"analyzedFiles\": 4, \"added\": 5, \"removed\": 13 },\n  \"findings\": [\n    { \"ruleId\": \"test-skipped\", \"severity\": \"high\", \"file\": \"test/cart.test.ts\", \"side\": \"added\", \"line\": 9,\n      \"message\": \"Test skipped with .skip\", \"snippet\": \"+ it.skip('applies SAVE10', () => {\" }\n  ]\n}\n```\n\nIt looks **only at added and removed lines** in the diff, so old code isn't re-flagged. Comments and string literals\nare lexed out before matching, which keeps `\".skip\"` inside a string or a comment from triggering anything. Lines that\nwere only *moved* (renames, file splits, reordering) are ignored. A test file that moved, or whose tests reappear in\nanother file, does not count as a deletion. Supported languages: **JavaScript/TypeScript, Python, Go, Rust, Java**\n(plus Kotlin test annotations), as well as CI and config files: GitHub Actions, GitLab CI, `package.json`, `tsconfig`,\njest/vitest/c8/nyc configs, `pyproject.toml`/` setup.cfg`/` tox.ini`/`.coveragerc`, Maven/Gradle, Makefiles, and\nhusky/shell hooks.\n\n| Rule | Default | What it catches | \n|---|---|---|\n| `test-file-deleted` | high | A test file was deleted (and not moved/renamed elsewhere in the diff). | \n| `test-count-dropped` | high | The net number of test cases (test()/it()/def test_/func Test/#[test]/@Test) went down. | \n| `test-skipped` | high | A skip was added: .skip, xit/xdescribe, @pytest.mark.skip/xfail, pytest.skip(), t.Skip(), #[ignore], @Disabled, @Ignore. | \n| `test-focused` | high | .only / fit / fdescribe silently disables every other test in the file or run. | \n| `test-conditional-skip` | low | A conditional skip (skipif, skipIf, importorskip, assumptions, `if testing.Short()` ) was added. | \n| `assertion-removed` | medium | Net assertion count in a test file dropped (expect/assert/t.Error/assert_eq!/assertEquals...). | \n| `assertion-weakened` | high | A specific assertion (toBe/toEqual/assert x == y/assertEquals) was replaced by a vague one (toBeTruthy/toBeDefined/assert x/assertNotNull). | \n| `assertion-trivial` | high | An assertion that can never fail was added (expect(true).toBe(true), assert True, assert!(true)). | \n| `assertion-expected-changed` | low | Only the literal expected value of an assertion changed. Confirm the code was wrong, not the test. | \n| `suppression-added` | medium | @ts-ignore, @ts-nocheck, @ts-expect-error, eslint-disable, # type: ignore, noqa, //nolint, #[allow(...)], @SuppressWarnings and friends. Blanket suppressions are medium, ones that name a specific rule are low, file/crate-wide ones are high. | \n| `coverage-exclusion-added` | low | istanbul/c8/v8 ignore, pragma: no cover, LCOV_EXCL, #[coverage(off)]. | \n| `ci-failure-ignored` | high | `\\|\\| true` , continue-on-error: true, allow_failure: true, --exit-zero, set +e on a test/lint/build command. | \n| `ci-step-removed` | high | A command that ran tests, lint or type checks was removed from CI config, scripts or package.json. | \n| `ci-step-disabled` | high | A CI job/step was disabled with `if: false` or`when: never` . | \n| `test-script-neutered` | high | The package.json test script was replaced with echo/true/exit 0. | \n| `test-exclusion-added` | medium | --passWithNoTests, -DskipTests, -x test, testPathIgnorePatterns, --ignore/--deselect, collect_ignore. | \n| `coverage-threshold-lowered` | high | A coverage threshold (coverageThreshold, fail_under, --cov-fail-under, thresholds, jacoco minimum...) was lowered or removed. | \n| `typecheck-weakened` | medium | tsconfig strict flags turned off, mypy ignore_errors / strict = false, pyright typeCheckingMode off. | \n| `lint-rule-disabled` | low | A lint rule was switched to \"off\"/0 in an ESLint config. | \n| `test-env-special-case` | high | Non-test source code branches on being under test (NODE_ENV === \"test\", JEST_WORKER_ID, \"pytest\" in sys.modules, testing.Testing(), cfg!(test)) or on CI. | \n| `error-swallowed` | medium | New empty catch / except: pass / .catch(() => {}) / if err != nil {}. | \n| `ignore-comment-added` | low | An inline fakegreen-ignore comment was added. Always reported so reviewers see what was waived. | \n\nSeverity is graded where it matters. For example, a blanket `# type: ignore` is medium, a targeted\n`# type: ignore[attr-defined]` is low, and a file-wide `// @ts-nocheck` is high. `|| true` on `npm test` is high, while\n`|| true` on `rm -rf build` is low. A conditional `skipif(sys.platform == \"win32\")` is low, but `skipif(True)` is high.\n\n**Skipped automatically:** Markdown/docs, lockfiles, vendored and `node_modules` code, generated files, and\n`fixtures/` / `testdata/` directories.\n\nClaude Code, Codex and Gemini CLI can run a command when the agent finishes its turn and **block** the stop, sending\nthe command's feedback back to the model. `fakegreen hook` speaks each agent's protocol. If it finds something at or\nabove `--fail-on`, the agent is told exactly what it faked and asked to fix it, or to stop and explain to you why the\nchange is intentional.\n\n```\nnpx fakegreen install claude     # .claude/settings.json   (--local → settings.local.json, --global → ~/.claude)\nnpx fakegreen install codex      # .codex/hooks.json       (--global → ~/.codex/hooks.json)\nnpx fakegreen install gemini     # .gemini/settings.json   (--global → ~/.gemini/settings.json)\n```\n\nThe installer shows a diff of the config change and asks before writing. Pass `--yes` to skip the prompt or\n`--dry-run` to only preview. It merges into existing config, is idempotent, and refuses to touch a file that isn't\nvalid JSON.\n\n## **Claude Code**: `.claude/settings.json`\n\n```\n{\n  \"hooks\": {\n    \"Stop\": [\n      {\n        \"hooks\": [\n          {\n            \"type\": \"command\",\n            \"command\": \"npx --yes fakegreen hook --agent claude\",\n            \"timeout\": 120,\n            \"statusMessage\": \"fakegreen: checking the diff for fake-green changes\"\n          }\n        ]\n      }\n    ]\n  }\n}\n```\n\nOn findings, the hook prints `{\"decision\":\"block\",\"reason\":\"...\"}` and Claude keeps working with the reason as its\nnext instruction. ([Claude Code hooks docs](https://code.claude.com/docs/en/hooks))\n\n## **Codex**: `.codex/hooks.json`\n\n```\n{\n  \"hooks\": {\n    \"Stop\": [\n      {\n        \"hooks\": [\n          {\n            \"type\": \"command\",\n            \"command\": \"npx --yes fakegreen hook --agent codex\",\n            \"timeout\": 120,\n            \"statusMessage\": \"fakegreen: checking the diff for fake-green changes\"\n          }\n        ]\n      }\n    ]\n  }\n}\n```\n\nOn findings, the hook prints `{\"decision\":\"block\",\"reason\":\"...\"}`, and Codex continues with the reason as a new\nprompt. Codex only loads project hooks when the project's `.codex/` layer is trusted, and it asks you to review and\ntrust each new hook (`/hooks`) before running it.\n([Codex hooks docs](https://developers.openai.com/codex/hooks))\n\n## **Gemini CLI**: `.gemini/settings.json`\n\n```\n{\n  \"hooks\": {\n    \"AfterAgent\": [\n      {\n        \"hooks\": [\n          {\n            \"name\": \"fakegreen\",\n            \"type\": \"command\",\n            \"command\": \"npx --yes fakegreen hook --agent gemini\",\n            \"timeout\": 120000,\n            \"description\": \"Block the turn when the diff fakes a green build\"\n          }\n        ]\n      }\n    ]\n  }\n}\n```\n\nOn findings, the hook prints `{\"decision\":\"deny\",\"reason\":\"...\"}`, which makes Gemini retry the turn with the reason\nas feedback. Gemini timeouts are in milliseconds. ([Gemini CLI hooks reference](https://geminicli.com/docs/hooks/reference/))\n\n**Cursor and Aider.** Neither has a blocking end-of-turn hook that fakegreen targets yet. Use the pre-commit hook and the\nagent skill below, or run `npx fakegreen --base main` before you accept the agent's work. Aider skips git hooks by\ndefault, so either enable them with `--git-commit-verify` or run `npx fakegreen --last-commit` after each Aider commit.\n\n**Hook details:**\n\n- **Clean diffs produce no output.** fakegreen exits 0 silently.\n- **Errors fail open.** If fakegreen itself breaks, your agent is never stuck.\n- **Loop guard.** If the agent was already blocked once (`stop_hook_active` ) and the findings haven't changed, the\nsecond stop is allowed and a warning is shown to you. That way an agent that has explained an intentional change\nisn't trapped.\n- **Choosing the diff.** By default the hook scans uncommitted work vs`HEAD` . If your agent commits as it goes, scan\nthe whole branch instead:`fakegreen install claude --hook-args \"--base origin/main\"` .\n- **Before the npm release, or with a local checkout:** point hooks at the build directly with`--command \"node /path/to/fakegreen/dist/cli.js\"` .\n\n```\nnpx fakegreen install pre-commit     # git pre-commit hook (respects core.hooksPath and .husky/pre-commit)\nnpx fakegreen install github-action  # writes .github/workflows/fakegreen.yml\nnpx fakegreen install skill          # copies SKILL.md to .claude/skills/fakegreen/ (--global → ~/.claude/skills)\n```\n\n**pre-commit framework** (`.pre-commit-config.yaml`):\n\n```\n- repo: https://github.com/fitzyracing1/fakegreen\n  rev: v0.1.0\n  hooks:\n    - id: fakegreen\n```\n\n**GitHub Actions.** Findings show up as annotations on the PR. See [`examples/github-workflow.yml`](https://github.com/fitzyracing1/fakegreen/blob/main/examples/github-workflow.yml):\n\n```\n- uses: actions/checkout@v4\n  with: { fetch-depth: 0 }\n- uses: fitzyracing1/fakegreen@v0.1.1   # composite action; inputs: version, base, fail-on, args\n  with:\n    fail-on: high\n```\n\n**Agent skill.** [`skills/fakegreen/SKILL.md`](https://github.com/fitzyracing1/fakegreen/blob/main/skills/fakegreen/SKILL.md) tells agents never to skip, delete or weaken\ntests to get green, never to special-case the test environment, and to run `npx fakegreen` before claiming they're\ndone. It works as a Claude Code skill. Paste it into `AGENTS.md`, `GEMINI.md`, `.cursor/rules` or `CONVENTIONS.md` for\nother agents.\n\nAdd an optional `.fakegreenrc.json`, a `.fakegreenrc`, or a `\"fakegreen\"` key in `package.json`:\n\n```\n{\n  \"rules\": {\n    \"coverage-exclusion-added\": \"off\",\n    \"suppression-added\": \"low\",\n    \"error-swallowed\": \"high\"\n  },\n  \"ignore\": [\"scripts/**\", \"**/generated/**\"],\n  \"testPatterns\": [\"e2e/**/*.ts\"],\n  \"failOn\": \"high\",\n  \"untracked\": true\n}\n```\n\n| Key | Meaning | \n|---|---|\n| `rules` | Per-rule `\"off\"` or a severity override (`\"high\"` ,`\"medium\"` ,`\"low\"` ). | \n| `ignore` | Globs for files to skip entirely. | \n| `testPatterns` | Extra globs for files that should be treated as tests. | \n| `failOn` | Default for `--fail-on` . | \n| `untracked` | Include untracked files in working-tree scans (default `true` ). | \n\nAdd a `fakegreen-ignore` comment on the line, or on the line above it. You can name rules, and adding a reason is a\ngood idea:\n\n```\n// fakegreen-ignore test-skipped -- flaky upstream API, tracked in #123\nit.skip('talks to the payments sandbox', async () => { ... });\nexcept Exception:  # fakegreen-ignore error-swallowed: best-effort telemetry\n    pass\n```\n\nFor removed lines (like a deleted assertion), put the comment anywhere among the added lines of the same hunk.\n`fakegreen-ignore-file` waives a whole file. Every ignore comment is itself reported as a **low** finding\n(`ignore-comment-added`), so a reviewer still sees what was waived. The hook's feedback explicitly tells agents not to\nadd waivers themselves. If you don't trust your agent with them at all, set\n`\"rules\": { \"ignore-comment-added\": \"high\" }` so any new waiver blocks the turn and gets surfaced to you.\n\n**Why not just ask an LLM to review the diff?**\nLLM reviewers are slow, cost money, need keys, and are easy to talk around. fakegreen runs in milliseconds, gives the\nsame answer every time, and works offline and in CI. Use both if you like. fakegreen is the cheap tripwire that runs on\nevery turn.\n\n**Will it flag my legitimate refactors?**\nSometimes, and that's the point of a review signal. Moves and renames are recognised, and so are tests that move\nbetween files and migrations from one CI command to another. In a dogfood run over the last 100 commits of 12 popular\nhuman-maintained repos (express, zod, vite, requests, flask, httpx, gin, cobra, ripgrep, clap, gson, spring-petclinic;\n1,200 commits), **34 high-severity findings** came up, about one every 35 commits. Almost all were real test removals\n(reverts, feature removals) that a reviewer would want to see anyway. Medium and low findings are context. The\ndefault `--fail-on high` only blocks on the high ones.\n\n**Does it understand my code (AST)?**\nNo. It uses a small lexer (strings and comments) plus targeted patterns on changed lines. That's what keeps it fast,\ndependency-free, and multi-language. See the limitations below.\n\n**Can the agent just disable fakegreen?**\nIt could edit the hook config, but that edit shows up in your diff. Combine the hook with the pre-commit hook or the\nGitHub Action, and treat changes to `.claude/`, `.codex/`, `.gemini/` or `.fakegreenrc.json` like changes to CI.\n\n**Does it send my code anywhere?**\nNo. It never makes a network call. It shells out to `git` and nothing else.\n\n- Heuristic, line-based analysis with no AST. Unusual formatting (a test declaration split across lines, regex literals that contain quotes, macros) can cause misses or occasional false positives.\n- Test-count compensation is diff-wide. If an agent deletes one test and adds an unrelated trivial one, the count doesn't drop, though a weakened or trivial assertion is still caught by the assertion rules.\n- Only JS/TS, Python, Go, Rust and Java/Kotlin source are analysed. Ruby, C#, C/C++, PHP, Swift and custom test DSLs are not, though their CI config changes still are.\n- The hook's default diff is uncommitted work vs `HEAD` . Use`--hook-args \"--base origin/main\"` if the agent commits\nduring the session.\n- `fakegreen-ignore-file` is only honoured when it appears in the diff's added lines or context.\n\nfakegreen is MIT-licensed and the CLI, hooks, pre-commit hook and GitHub Action will stay free. I'm Joshua Almeida, and I build and maintain it on my own. If it saves you from merging a \"fixed\" test suite that was really a deleted one, here are three ways to help:\n\n**Sponsor the project.** [GitHub Sponsors](https://github.com/sponsors/fitzyracing1) pays for the time I spend on new\nrules, new languages and false-positive fixes. Every sponsorship helps, small ones included.\n\n**fakegreen for Teams (early access waitlist).** I'm looking into a hosted version for teams running AI coding agents\nacross many repos. The plan is a GitHub App that posts findings as PR comments and check runs, one shared policy for the\nwhole org, and a history of fake-green incidents grouped by agent and repo. **Nothing is built yet.** If your team would\nuse it, [join the waitlist](https://fitzyracing1.github.io/fakegreen/#teams) and tell me what you'd need, because that\ndecides what gets built.\n\n**Consulting & custom rules.** If your team is rolling out Claude Code, Codex, Gemini CLI, Cursor or Aider, I can\nhelp you write fakegreen rules for your own codebase, test conventions and CI setup, and wire up guardrails (hooks,\npre-commit, CI gates) that agents can't quietly route around. Email\n[fitzyracing1@gmail.com](mailto:fitzyracing1@gmail.com?subject=fakegreen%20consulting).\n\nSee [CONTRIBUTING.md](https://github.com/fitzyracing1/fakegreen/blob/main/CONTRIBUTING.md). New detectors need a true-positive fixture **and** a false-positive guard.\n\n[MIT](https://github.com/fitzyracing1/fakegreen/blob/main/LICENSE) © 2026 Joshua Almeida", "url": "https://wpnews.pro/news/fakegreen-catch-ai-coding-agents-faking-a-green-build-no-llm", "canonical_source": "https://github.com/fitzyracing1/fakegreen", "published_at": "2026-10-08 15:27:24+00:00", "updated_at": "2026-10-08 15:48:05.317598+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-tools"], "entities": ["fakegreen", "Claude Code", "Codex", "Cursor", "Gemini CLI", "Aider", "Node.js", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fakegreen-catch-ai-coding-agents-faking-a-green-build-no-llm", "markdown": "https://wpnews.pro/news/fakegreen-catch-ai-coding-agents-faking-a-green-build-no-llm.md", "text": "https://wpnews.pro/news/fakegreen-catch-ai-coding-agents-faking-a-green-build-no-llm.txt", "jsonld": "https://wpnews.pro/news/fakegreen-catch-ai-coding-agents-faking-a-green-build-no-llm.jsonld"}}