{"slug": "how-i-screen-an-ai-coding-agent-before-letting-it-near-my-repo", "title": "How I Screen an AI Coding Agent Before Letting It Near My Repo", "summary": "A developer outlined a six-point trial for evaluating autonomous AI coding agents before granting them access to production repositories, centered on a single scoped task such as adding input validation to a known function. The checklist covers plan-before-edit behavior, permission scoping, blast radius, failure behavior, context strategy, and cost per completed task, with the recommendation to treat agents as fast junior developers whose output still requires review.", "body_md": "Every week another \"autonomous AI coding agent\" launches, and every week someone I know lets it loose on a production repository, then spends the evening reviewing a 900-line diff that touches files nobody asked it to touch.\n\nThe fix isn't a better benchmark table. It's a cheap trial.\n\nPick one function you know well - ideally slightly messy, with a couple of edge cases - and give the agent exactly one instruction:\n\n\"Add input validation to `parseConfig` and a test that covers the empty-string and null cases. Don't change anything else.\"\n\nThen watch four things:\n\n**1. Plan before edit.** You want to see the intended change list before anything is written to disk.\n\n**2. Permission scoping.** Can it run shell commands? Which ones? Can it read `.env` files, SSH keys, or your cloud credentials? A sandbox or an approval prompt isn't friction - it's the whole safety model.\n\n**3. Blast radius.** How many files does a typical task touch? An agent that \"helpfully\" reformats your project while fixing a bug is worse than useless in a team.\n\n**4. Failure behaviour.** Retry loops that quietly change the goal are the single most expensive failure mode. You want loud, early, specific failure.\n\n**5. Context strategy.** Does it index the repository, or only see the files you mention? This decides whether it will follow your existing patterns or happily introduce a second HTTP client.\n\n**6. Cost per completed task.** Token pricing tells you almost nothing. Divide your monthly spend by merged pull requests.\n\nAnything that passes the trial goes on a short list with a note about what it's good at: refactors, tests, glue code, migrations. I keep the longer teaching version of this checklist - the one with the exact prompts - at [haiai123's beginner guide to choosing an AI coding agent](https://www.haiai123.com/en/blog/how-to-choose-an-ai-coding-agent-beginner-checklist), and I check the [coding leaderboard](https://www.haiai123.com/en/rankings/coding) before any of it to see how the underlying models are moving.\n\nTreat the agent as a very fast junior developer who has never read your codebase conventions and forgets everything overnight. You wouldn't merge a junior's branch without reading it, and you wouldn't let one run `rm -rf` unsupervised. The agent changes the speed of writing code - not the standard for shipping it.", "url": "https://wpnews.pro/news/how-i-screen-an-ai-coding-agent-before-letting-it-near-my-repo", "canonical_source": "https://dev.to/_8def5737f8730de95bc297/how-i-screen-an-ai-coding-agent-before-letting-it-near-my-repo-40h3", "published_at": "2026-10-01 02:34:30+00:00", "updated_at": "2026-10-01 02:46:35.935629+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "ai-products"], "entities": ["haiai123"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-i-screen-an-ai-coding-agent-before-letting-it-near-my-repo", "markdown": "https://wpnews.pro/news/how-i-screen-an-ai-coding-agent-before-letting-it-near-my-repo.md", "text": "https://wpnews.pro/news/how-i-screen-an-ai-coding-agent-before-letting-it-near-my-repo.txt", "jsonld": "https://wpnews.pro/news/how-i-screen-an-ai-coding-agent-before-letting-it-near-my-repo.jsonld"}}