I Built an Open-Source QA Agent Because AI Shouldn’t Grade Its Own Homework A developer built MaruCheck, an open-source QA and verification tool that creates an independent verification layer between AI coding agents and shipped software. The tool uses a 'Quality Contract' to maintain a persistent representation of intended product behavior, detecting semantic changes rather than silently rewriting tests to match new implementations. It also analyzes diffs to assess risk and prioritize verification for critical code changes. AI coding tools are getting ridiculously good. You can describe a feature, let an agent implement it, ask it to write tests, and sometimes have something working in minutes. But there’s a problem I kept thinking about: What happens when the same AI writes the code and then validates its own understanding of the requirement? If the model misunderstands the requirement once, it can easily carry that same misunderstanding into the tests. You end up with something like this: Product requirement ↓ Coding agent ↓ Incorrect interpretation ↓ Implementation ↓ Same agent writes tests ↓ Tests validate the same interpretation ↓ Everything passes The test suite is green. The product is still wrong. That idea eventually became MaruCheck . MaruCheck is an open-source QA and verification tool designed around AI-generated software. The goal is not to replace Playwright, Vitest, Jest, CI, or existing testing infrastructure. The goal is to create an independent verification layer between the coding agent and the software being shipped. The basic idea looks like this: PRODUCT INTENT │ ┌───────┴───────┐ ↓ ↓ CODING AGENT MARUCHECK ↓ ↓ CODE QUALITY CONTRACT │ │ └───────┬───────┘ ↓ VERIFY ↓ PASS / BLOCK The coding agent builds. MaruCheck verifies. One of the main ideas behind MaruCheck is something I call a Quality Contract . Instead of generating tests entirely from the implementation, the system keeps a persistent representation of what the feature is actually supposed to do. A simplified example: feature: subscription-management requirements: - Free users get 10 generations per month - Pro users have unlimited generations - Successful upgrades take effect immediately - Cancellation keeps Pro access until period end invariants: - Failed payments must never activate Pro - Client-controlled data cannot prove payment - Replayed webhooks must be idempotent Now imagine an AI agent changes the subscription implementation. MaruCheck can compare that change against the expected behavior instead of assuming that the new implementation defines the truth. This is one of the features I care about most. Suppose the approved behavior says: Free users can upload 5 files. Then an AI-generated change does this: js - const FREE LIMIT = 5 + const FREE LIMIT = 10 A test-generation system could potentially see the new implementation, regenerate the test, and now expect 10. Everything stays green. MaruCheck should not do that. Instead, it treats the mismatch as a semantic change: SEMANTIC CHANGE DETECTED Expected: 5 Observed: 10 Was this intentional? The system can propose changing the contract, but it should not silently rewrite product behavior just to make the tests pass. That distinction matters. Another thing I didn’t want was a tool that blindly runs every possible test after every change. Not every diff has the same risk. Changing a marketing page is not the same as changing: billing-webhook.ts authentication.ts permissions.ts subscription-state.ts So MaruCheck analyzes the diff and tries to understand what is actually affected. For example: Changed: billing-webhook.ts Risk: HIGH Why: + payment-sensitive code changed + subscription state transition changed + related historical regression exists + critical Quality Contract affected That risk assessment can then influence what verification gets run. This became another part of the project that I really liked. Imagine six months ago your application had this bug: BUG 143 A user could access another user's invoice by changing the invoice ID. The bug gets fixed. Everyone moves on. Months later, a coding agent modifies invoice authorization again. A normal testing tool mostly sees the current code. MaruCheck can remember: This area previously caused an authorization regression. and bring the relevant regression checks back into the verification plan: ✓ cross-account invoice checks ✓ IDOR verification ✓ ownership tests ✓ authorization boundary tests The idea is that QA knowledge should accumulate over time instead of disappearing into old tickets and forgotten incidents. Some of the functionality built so far includes: The project is designed to work with existing testing tools rather than trying to reinvent all of them. The interesting part, to me, is the reasoning and orchestration layer around those tools. I also wanted MaruCheck to fit directly into the workflow developers already have with coding agents. So an agent like Codex, Claude Code, Cursor, or another MCP-compatible tool can eventually follow a workflow like: Implementation complete. maru verify feature "subscription cancellation" MaruCheck performs an independent verification. If it finds something: BLOCKING ISSUE SUB-004 violated. Expected: Pro access remains active until period end. Actual: Account becomes FREE immediately. The coding agent can fix the implementation and ask MaruCheck to verify it again. That creates a loop where the builder and verifier have separate responsibilities. The name is inspired by the Kobayashi Maru idea. Not because I wanted to make a Star Trek-themed testing tool, but because I liked the underlying concept: Don’t only test what happens when everything goes according to plan. Test what happens when assumptions break. That maps nicely to the kind of QA I want MaruCheck to do: What if the webhook fires twice? What if the payment succeeds but the database update fails? What if two requests happen at the same time? What if permissions change during the workflow? What if an old regression comes back? What if the implementation passes its tests but still violates the original product requirement? Hence: MaruCheck — Test what your AI didn’t. I initially thought about keeping the project more closed, but I eventually decided that open source made much more sense. This kind of developer infrastructure gets better when developers can: I’m especially interested in seeing how it behaves against real projects being built heavily with AI coding agents. Website: Core repository: https://github.com/Kidus-M/MaruCheck https://github.com/Kidus-M/MaruCheck Website repository: https://github.com/Kidus-M/MaruCheck-Web https://github.com/Kidus-M/MaruCheck-Web If the idea sounds interesting, I’d really appreciate: But the feedback I want most is technical criticism. Does the independent-verifier model make sense? Would you actually keep Quality Contracts inside your repository? What would a tool like this have to catch before you trusted it enough to run on every pull request? And most importantly: What am I getting wrong? I’d love to hear your thoughts.