Defending your codebase from AI slop: guardrails that keep you in control A developer outlined a guardrail stack for keeping AI coding agents from degrading a codebase, combining ARCHITECTURE.md and AGENTS.md convention files with dependency-cruiser for enforcing allowed import directions between layers and Semgrep for pattern-based rules that flag misplaced helpers and other structural violations. The approach writes standards in a machine-enforceable form so checks fail the build with messages that explain why, rather than relying on human review to catch the volume of agent-generated pull requests. AI coding agents are fast. They also write code that compiles, looks fine at a glance, and slowly wears down your architecture. A component imports the database layer directly. A function grows to 200 lines with three nested ternaries. A dependency shows up that nobody asked for. A routes file picks up a formatDate helper, then a slugify , then a retryWithBackoff . Tests run every line but never check a result. None of this is new. Junior developers in a hurry do the same things. What's new is the volume. If an agent opens ten PRs a day, careful human review alone won't catch everything. The answer isn't to stop using AI. It's to write your standards down in a form machines can enforce, so the rules hold no matter who or what wrote the code. Here's the stack I usually build for my projects. Tools can only enforce rules that exist. Before adding any tooling, I decide how I want my code to look, define the rules, and write them in two files: ARCHITECTURE.md covers the layers, which layer may import which, where state lives, and how data flows. AGENTS.md , essential nowadays, is usually created automatically and then edited by me. It covers conventions for AI agents: naming, patterns to use, patterns to avoid, and tools that are off-limits for example: "we use Biome, don't add ESLint" . Agents read these files. So do new hires. And every tool below points back to a section in one of them, so when a check fails, the message explains why and not just what . Something I've discovered quite recently: dependency-cruiser https://www.npmjs.com/package/dependency-cruiser checks which files import which. You describe the allowed directions between layers, and it fails the build when something crosses a line: This is the kind of mistake AI makes most, because it takes whatever import path is closest. A rule that always gives the same answer stops it before review. dependency-cruiser tells you which files talk to each other . It can't tell you what the code inside a file is doing . That's what Semgrep https://semgrep.dev is for. Semgrep is a pattern matcher that understands code. You write a rule that says "code shaped like X, in files matching Y, is a violation." Semgrep parses each file into a syntax tree and reports every match. It doesn't run your code or try to understand your whole app. It only checks shapes. A rule is a short YAML entry: rules: - id: api-routes-no-helpers languages: python severity: ERROR message: routes.py is for route handlers only. Move helpers to their own module see ARCHITECTURE.md → API . paths: include: packages/api/src/api/routes.py patterns: - pattern: | def $FUNC ... : ... - pattern-not-inside: | @$ROUTER.$METHOD ... def $F ... : ... pattern is the code to match. $FUNC matches any name, and ... matches anything. pattern-not-inside is the exception. Here, functions with a route decorator are allowed. paths limits the rule to certain files. message is what the developer or the AI agent sees when the rule fails. Write it as an instruction. Running semgrep --config .semgrep/rules/ checks the whole repo against every rule in that folder. We keep one YAML file per area app, API, mq so rules stay easy to find, and each rule gets a small test fixture with code that should match and code that shouldn't, so we know the rule catches what we meant. Read their docs https://docs.semgrep.dev/writing-rules/generic-pattern-matching or ask your agent to on what's possible and get creative. This was the rule we wanted most. Some files exist for one job, and AI agents love to drop "just one small helper" into them. After a few months, your routes file is half utilities. So we made it impossible. Each of these files may only contain its one kind of thing: routes.py : route handlers only services/ functions/ Helpers go in their own module, where they can be named, tested and reused. The rule message tells the agent exactly that, so it usually fixes itself on the next try. Once you think in shapes, a lot of review comments turn into rules: fetch in routes and server functions db/queries tx: DbClient = db this hides transaction bugs Semgrep works on many languages https://docs.semgrep.dev/supported-languages , so it covers anything you want. Rule of thumb: if a review comment can be written as a code pattern, make it a Semgrep rule. Then it gets checked the same way on every PR. You probably have violations already. Don't let that stop you. In CI, Semgrep can compare against a baseline: main New code is held to the standard from day one, and nobody has to fix everything first. Architecture rules keep the code tidy. They don't tell you whether it's safe. AI agents write the same security bugs humans do, just faster: SQL built from strings, an endpoint that forgets to check who's calling it, a token pasted into a config file, user input rendered as HTML. Security tools fall into three groups, and each sees something the others can't: | Kind | Looks at | Finds | Misses | |---|---|---|---| | SAST static application security | Your source code, not running | Injection, unsafe APIs, hardcoded secrets, tainted data flows | Runtime config, auth bugs that depend on real data | | SCA software composition analysis | Your dependencies and lockfiles | Known CVEs, malicious packages, licenses | Bugs in your own code | | DAST dynamic application security | The running app, from outside | Missing auth, exposed endpoints, bad headers, misconfigured servers | Where in the code the bug is; code paths it never reaches | SAST and SCA run on every PR in seconds or minutes. DAST needs a deployed app, so it runs against a preview environment or staging. Most of the tools below cover more than one group: | Tool | SAST | SCA | DAST | Other | |---|---|---|---|---| | Semgrep | ✅ | ✅ | | Secrets | | Trivy | | ✅ | | Containers, IaC, secrets | | Snyk | ✅ | ✅ | ✅ | Containers, IaC | | Checkmarx One | ✅ | ✅ | ✅ | Containers, IaC, secrets, API | | GitHub CodeQL | ✅ | | | Dependabot covers SCA | | ZAP | | | ✅ | Free, open source | I am an OSS supporter, so I usually go with Semgrep and Trivy. Companies I've worked for went with whatever suited them best. Choose any; no judgement there. Good SAST tools track data flow: they follow user input from a request handler and flag it if it reaches SQL, a shell or HTML unescaped. Ask an agent to "add search" and you may get: python @router.get "/search" def search q: str, db: Session = Depends get db : return db.execute text f"SELECT FROM products WHERE name LIKE '%{q}%'" It works, and it's SQL injection. Semgrep Code https://semgrep.dev/products/semgrep-code catches it with the engine you already run for architecture add --config p/owasp-top-ten . CodeQL https://codeql.github.com is free for public repos and shows results in the PR. Snyk Code https://snyk.io/product/snyk-code/ and Checkmarx https://checkmarx.com cover the same ground in their platforms. Fail the build on high-confidence rules only. A noisy check gets switched off. Agents add packages freely, sometimes outdated, sometimes made up and squatted "slopsquatting" . Trivy https://trivy.dev is the free baseline for vulnerable packages, secrets, container images and IaC: trivy fs . --scanners vuln,secret,misconfig --severity HIGH,CRITICAL --exit-code 1 Trivy only matches versions, so it flags vulnerable packages whether or not you use the vulnerable part. Semgrep Supply Chain https://semgrep.dev/products/semgrep-supply-chain checks reachability: does your code actually call the vulnerable function? Snyk https://snyk.io opens fix PRs and alerts when a new CVE hits code you already shipped. Checkmarx One https://checkmarx.com also detects malicious packages such as typosquats. Require CODEOWNERS approval on package.json and lockfiles, so every new dependency is a human decision. DAST sends real requests to a deployed app, so it catches what code scanning can't: an endpoint the agent left without an auth check, one user reading another's data by changing an ID, missing security headers, exposed debug pages. Run it against a preview deploy or staging: docker run -t ghcr.io/zaproxy/zaproxy:stable zap-api-scan.py \ -t https://preview.example.com/openapi.json -f openapi ZAP https://www.zaproxy.org is free. Nuclei https://github.com/projectdiscovery/nuclei checks for known exposures. StackHawk https://www.stackhawk.com , Snyk API & Web https://snyk.io/product/snyk-api-web/ and Checkmarx DAST https://checkmarx.com/product/dast/ are paid options. Give the scanner credentials and an API spec, or it will scan your login page and little else. DAST, I'd say, is a nice-to-have, not a must-have. Folks at Semgrep think so too https://semgrep.dev/blog/2023/dast-devsecops/ . Occasional manual testing or pentesting is more than enough. Something I've learned about from Uncle Bob https://x.com/unclebobmartin/status/2047661738456121506?s=20 . I never thought about it until AI became a thing. CRAP stands for Change Risk Anti-Patterns, a metric used to identify risky, complex, and poorly tested code. The CRAP Index measures the maintenance risk of a specific function or method. Rule of thumb: a score above 30 indicates a high-risk, "CRAPpy" method that needs attention or refactoring. I found two tools I could integrate into my pipelines that would help to keep AI-generated code under control: Some rules can't be written as patterns: "does this name match what the code does?" or "is this abstraction premature?" For those, we can use CodeRabbit https://docs.coderabbit.ai/triage/rules or one of many other AI code review tools with a strict config file, like .coderabbit.yaml : The important part is that CodeRabbit is the last layer, not the first. AI review isn't consistent, so anything that can be checked by a tool that always gives the same answer should be. CodeRabbit handles judgement calls. What we've set up covers most of it. Here are a few more tools that can help you keep even tighter control over your repo: Every layer follows the same idea: move each rule to the most reliable tool that can enforce it. AI can write as much code as it likes. It just has to pass the same checks as everyone else. That's how you stay in control.