cd /news/ai-tools/agentcheck-regression-testing-for-ai… · home topics ai-tools article
[ARTICLE · art-106489] src=github.com ↗ pub= topic=ai-tools verified=true sentiment=· neutral

AgentCheck – regression testing for AI agents, with diff-aware CI reports

AgentCheck, a new open-source regression testing tool for AI agents, lets developers define expected agent behavior in YAML and run pass/fail checks against any CLI command or HTTP endpoint, with an LLM judge scoring responses. It integrates into CI via GitHub Actions, supports diff-aware reports against a baseline, and can post PR comments summarizing regressions and improvements. The tool targets a gap left by observability platforms like Langfuse, Braintrust, and Arize, which focus on production monitoring rather than lightweight pre-deployment checks.

read4 min views1 publishedAug 21, 2026
AgentCheck – regression testing for AI agents, with diff-aware CI reports
Image: Michielbdejong (auto-discovered)

Regression testing for AI agents. Define what your agent should do in plain YAML, run it against your actual agent (any CLI command or HTTP endpoint), and let an LLM judge score every response against your stated criteria — pass/fail, with a reason. Wire it into CI so a prompt change, a tool swap, or a model upgrade can't silently break behavior your users depend on.

This is deliberately narrow: it is not a production-observability platform (that's Langfuse / Braintrust / Arize territory, and they're well-funded — don't compete head-on). It's the thing almost nobody has built well yet: a fast, dev-friendly pre-deployment check that fits in a GitHub Actions step the same way pytest

does.

  • Full observability/eval platforms have raised $50–80M rounds in the last year and are actively crowding the "log + monitor production agent traffic" space.
  • Almost none of them are built as a lightweight, git-native regression suite a solo developer can add to CI in five minutes — that gap is the wedge.
  • Distribution is self-serve/PLG (open-source CLI, developer audience) rather than enterprise sales — the thing solo AI founders are consistently worst at.
pip install -e .
export ANTHROPIC_API_KEY=sk-...
agentcheck run examples/tests.yaml

See examples/tests.yaml

. Each test case specifies an input, a plain-English description of what a correct response looks like, and one of two ways to reach your agent:

command: "..."

— run it as a subprocess;input

is piped to stdin, stdout is captured as the output. Works with any language.agent: "module.path:function_name"

— import that module and call the function in-process withinput

as its only argument; its return value is the output. Useful for LangGraph/CrewAI/Claude-Agent-SDK-style agents that are Python callables rather than standalone CLI scripts — seeexamples/inprocess_agent.py

.

Exactly one of the two is required per case. Either way, the output is scored pass/fail with a one-line reason by an LLM judge — no brittle string matching.

agentcheck run tests.yaml --json-out results.json

writes a JSON report you can upload as a build artifact (see examples/.github/workflows/agentcheck.yml

).

Add --post-pr-comment

and, on a pull-request run with GITHUB_TOKEN

set (the job needs permissions: pull-requests: write

), agentcheck posts a markdown summary table as a PR comment, updating the same comment on repeat runs instead of piling up new ones. It's a silent no-op everywhere else (pushes, local runs), so it's safe to leave on in every CI invocation.

A flat pass count ("18/20 passed") doesn't tell you whether a change helped or hurt — you have to go read the table. --baseline

fixes that by diffing the current run against a previous --json-out

report, keyed by test name:

agentcheck run tests.yaml --json-out results.json --baseline baseline.json

Try it locally against the bundled example:

agentcheck run examples/tests.yaml --baseline examples/baseline.json

Every test lands in one bucket: unchanged (same pass/fail as the baseline), regressed (baseline passed, now fails — this is the one you care about), improved (baseline failed, now passes), new (not in the baseline), or removed (in the baseline but not in this run — probably a deleted test case, worth a glance). The console prints a one-line summary plus a table of regressions and improvements; a missing or unreadable baseline (there's no baseline yet on a repo's first run) prints a warning and falls back to the plain pass/fail report instead of failing the whole run.

--post-pr-comment

picks this up automatically when --baseline

is also set, so the PR comment leads with "vs baseline: 2 unchanged, 1 improved, 1 regressed" and calls out the regressions specifically, instead of just restating the full results table.

To actually wire this into CI you need somewhere for the baseline to come from — the usual pattern is: on every push to your default branch, run agentcheck with --json-out results/baseline.json

and commit that file back to the repo; on every PR, read the base branch's copy of that file (git show origin/main:results/baseline.json

) and pass it as --baseline

. See examples/.github/workflows/agentcheck.yml

for a full working version of that.

pip install -e ".[dev]"
pytest

This week: get this CLI working end to end against your own toy agent. Dogfood it.Week 2: open-source it. Post it where agent builders actually hang out (r/LocalLLaMA, the LangChain/LlamaIndex Discords, Hacker News "Show HN", relevant X threads). The goal isn't virality — it's finding 10-20 people who hit this exact pain point and will tell you what's missing.Week 3-4: add the thing they ask for most. Likely candidates: a hosted dashboard for run history, Slack/GitHub PR-comment reporting, or support for a specific popular agent framework (LangGraph, CrewAI, the Claude Agent SDK) as a first-class integration.Only once people are using the free CLI regularly: introduce a paid hosted tier (run history, team sharing, trend charts) — don't build monetization before you have free users who'd miss it if it disappeared.

Don't try to become a general observability platform — that's the crowded, well-funded lane. Stay the "fast, git-native regression check" tool. Depth in one narrow job beats breadth against funded competitors.

MIT — see LICENSE

.

── more in #ai-tools 4 stories · sorted by recency
── more on @agentcheck 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agentcheck-regressio…] indexed:0 read:4min 2026-08-21 ·