# AI Coding Tools Are Great Until the Only Test Is a Device in Someone's Hand

> Source: <https://dev.to/nabeelbaghoor/ai-coding-tools-are-great-until-the-only-test-is-a-device-in-someones-hand-3apk>
> Published: 2026-10-08 21:35:03+00:00

Most "how I use AI coding tools" posts come from someone who lives in one web framework all day. That is a fair place to form an opinion, but it hides the thing I find most interesting: these tools are not equally good everywhere, and the quality drops in a very predictable direction.

My week is less tidy than one framework. In the same few days I might touch a Unity project for the RAQTS sports-tech platform, a React/TypeScript/GraphQL dashboard for Bettershop, and an n8n workflow behind a Retell voice agent for Tested Media's CallSetter AI. Before that came a C# SDK running on Windows, Android and iOS at Geonode, and AR work at ARCortex. I use Claude Code and Cursor across all of it, every working day.

Working across that spread taught me one rule that holds up better than "always use it" or "never trust it":

**AI coding tools are excellent where there is a lot of consistent public code and a fast way to check the result. They get worse as you move toward niche platform APIs, editor-owned file formats, and systems where the only real test is a device in someone's hand.**

So the answer changes by layer. Here is how it splits for me.

The clearest win I have had was a pipeline for building and testing voice agents at Fortell AI, which I wrote up in [How I use Claude Code and Comet to build and test AI voice agents in a day](https://dev.to/nabeelbaghoor/how-i-use-claude-code-and-comet-to-build-and-test-ai-voice-agents-in-a-day-3ikl). The important part was not the prompt. It was the spec. Once the fields to capture, the tools the agent can call and the escalation rules were written down plainly, Claude Code could generate the function definitions, the n8n webhook skeletons and the test scenarios in one pass.

That generalises. A vague request gets a plausible-looking guess. A precise one gets something close to what I would have typed, faster. Most of the skill is in the spec.

Anything declarative and checked into git is a good fit: workflow definitions, CI YAML, environment templates, types generated from a GraphQL schema, voice agent config exported out of a vendor dashboard. These have a clear shape, a clear diff and usually a validator. Config in files is config a tool can read, change and explain, and config a reviewer can diff. Config that only exists in a web dashboard is invisible to both.

As a fractional CTO I regularly walk into code someone else built. The first days used to go on tracing data flows by hand. Now I ask the tool to map them: where a value is written, which screens read it, what calls this endpoint.

I still verify the map against the code, because a confident wrong answer about architecture is worse than no answer. But as a way to get oriented it is the biggest time saver I have found, and the least risky use, because nothing it says ships.

Unit tests on pure logic, fixtures, edge-case inputs, frozen regression scenarios for a voice agent. A wrong test usually fails loudly, which makes tests a forgiving thing to generate. The real benefit is economic: writing the boring set becomes cheap enough that it actually gets written.

This is where I have seen the most confident nonsense. On a React Native VPN app, the real behaviour lives in Android's `VpnService` and iOS Network Extensions, with entitlements, separate processes and background rules that differ per OS and per version (I went into that split in [A React Native VPN Is Two Programs, and Only One of Them Is JavaScript](https://dev.to/nabeelbaghoor/a-react-native-vpn-is-two-programs-and-only-one-of-them-is-javascript-475j)).

AI tools will happily produce a method that does not exist, a permission flow that was deprecated two versions ago, or a lifecycle assumption that only fails on a real device after the app has been backgrounded for ten minutes. None of those fail at compile time in an obvious way. Some of them do not fail until a user files a ticket.

I still use the tool here, but as a fast reader of documentation that I then check myself, never as the author of code I accept on trust. The same went for the cross-platform C# networking core at Geonode: the shared logic was fine to generate against, the per-platform edges were not.

Unity serialises scenes and prefabs as YAML, which looks exactly like text an AI can edit. In practice those files are full of object IDs and references the editor manages, and a hand edit that is syntactically fine can quietly break a reference you only notice at runtime.

So the line I draw is about ownership of the file format, not about any one model. I let the tool write C# scripts, editor tooling and build scripts. I do not let it edit what the Unity editor owns. If a human should not be hand-editing a file, neither should a model.

In AR, "does it work" often means "does it line up with the real world when you stand there holding the phone". No coding tool can see that. The coordinate maths can be generated, but whether an anchor drifts half a metre on a real device is a field test, not a code review.

Some things I do not delegate, however good the tools get:

The honest version is that AI tools make me faster on the parts of a project that were never the hard part: wiring, boilerplate, config, test scaffolding, reading unfamiliar code. That time goes back into what actually decides whether the product works, which is architecture, edge cases, platform-specific behaviour and conversations about what the client really needs.

What they do not replace is the person who knows why a Unity reference broke, why an iOS extension got killed in the background, or why a voice agent should transfer a call instead of answering it. If anything they raise the value of that knowledge, because they produce a lot of plausible output and someone has to know which parts are wrong.

The rule underneath all of it is short: the tool writes, I own.

I am curious where other people have drawn their line. If you work on a stack that is not mostly web, which layer did you stop trusting first?

*Originally published on [nabeelbaghoor.com](https://nabeelbaghoor.com/blog/claude-code-on-client-projects/).*
