Every week another "autonomous AI coding agent" launches, and every week someone I know lets it loose on a production repository, then spends the evening reviewing a 900-line diff that touches files nobody asked it to touch.
The fix isn't a better benchmark table. It's a cheap trial.
Pick one function you know well - ideally slightly messy, with a couple of edge cases - and give the agent exactly one instruction:
"Add input validation to parseConfig and a test that covers the empty-string and null cases. Don't change anything else."
Then watch four things:
1. Plan before edit. You want to see the intended change list before anything is written to disk.
2. Permission scoping. Can it run shell commands? Which ones? Can it read .env files, SSH keys, or your cloud credentials? A sandbox or an approval prompt isn't friction - it's the whole safety model.
3. Blast radius. How many files does a typical task touch? An agent that "helpfully" reformats your project while fixing a bug is worse than useless in a team.
4. Failure behaviour. Retry loops that quietly change the goal are the single most expensive failure mode. You want loud, early, specific failure.
5. Context strategy. Does it index the repository, or only see the files you mention? This decides whether it will follow your existing patterns or happily introduce a second HTTP client.
6. Cost per completed task. Token pricing tells you almost nothing. Divide your monthly spend by merged pull requests.
Anything that passes the trial goes on a short list with a note about what it's good at: refactors, tests, glue code, migrations. I keep the longer teaching version of this checklist - the one with the exact prompts - at haiai123's beginner guide to choosing an AI coding agent, and I check the coding leaderboard before any of it to see how the underlying models are moving.
Treat the agent as a very fast junior developer who has never read your codebase conventions and forgets everything overnight. You wouldn't merge a junior's branch without reading it, and you wouldn't let one run rm -rf unsupervised. The agent changes the speed of writing code - not the standard for shipping it.