{"slug": "ai-coding-tools-are-great-until-the-only-test-is-a-device-in-someone-s-hand", "title": "AI Coding Tools Are Great Until the Only Test Is a Device in Someone's Hand", "summary": "A developer who works across Unity, React/TypeScript, n8n and native mobile projects reports that AI coding tools like Claude Code and Cursor perform best where consistent public code and fast verification exist, and degrade sharply on niche platform APIs, editor-owned file formats, and systems whose only real test is a physical device. The account cites a Fortell AI voice-agent pipeline built in a day from a written spec, and a React Native VPN app where the tools generated non-existent methods and deprecated permission flows that failed only on real hardware.", "body_md": "Most \"how I use AI coding tools\" posts come from someone who lives in one web framework all day. That is a fair place to form an opinion, but it hides the thing I find most interesting: these tools are not equally good everywhere, and the quality drops in a very predictable direction.\n\nMy week is less tidy than one framework. In the same few days I might touch a Unity project for the RAQTS sports-tech platform, a React/TypeScript/GraphQL dashboard for Bettershop, and an n8n workflow behind a Retell voice agent for Tested Media's CallSetter AI. Before that came a C# SDK running on Windows, Android and iOS at Geonode, and AR work at ARCortex. I use Claude Code and Cursor across all of it, every working day.\n\nWorking across that spread taught me one rule that holds up better than \"always use it\" or \"never trust it\":\n\n**AI coding tools are excellent where there is a lot of consistent public code and a fast way to check the result. They get worse as you move toward niche platform APIs, editor-owned file formats, and systems where the only real test is a device in someone's hand.**\n\nSo the answer changes by layer. Here is how it splits for me.\n\nThe clearest win I have had was a pipeline for building and testing voice agents at Fortell AI, which I wrote up in [How I use Claude Code and Comet to build and test AI voice agents in a day](https://dev.to/nabeelbaghoor/how-i-use-claude-code-and-comet-to-build-and-test-ai-voice-agents-in-a-day-3ikl). The important part was not the prompt. It was the spec. Once the fields to capture, the tools the agent can call and the escalation rules were written down plainly, Claude Code could generate the function definitions, the n8n webhook skeletons and the test scenarios in one pass.\n\nThat generalises. A vague request gets a plausible-looking guess. A precise one gets something close to what I would have typed, faster. Most of the skill is in the spec.\n\nAnything declarative and checked into git is a good fit: workflow definitions, CI YAML, environment templates, types generated from a GraphQL schema, voice agent config exported out of a vendor dashboard. These have a clear shape, a clear diff and usually a validator. Config in files is config a tool can read, change and explain, and config a reviewer can diff. Config that only exists in a web dashboard is invisible to both.\n\nAs a fractional CTO I regularly walk into code someone else built. The first days used to go on tracing data flows by hand. Now I ask the tool to map them: where a value is written, which screens read it, what calls this endpoint.\n\nI still verify the map against the code, because a confident wrong answer about architecture is worse than no answer. But as a way to get oriented it is the biggest time saver I have found, and the least risky use, because nothing it says ships.\n\nUnit tests on pure logic, fixtures, edge-case inputs, frozen regression scenarios for a voice agent. A wrong test usually fails loudly, which makes tests a forgiving thing to generate. The real benefit is economic: writing the boring set becomes cheap enough that it actually gets written.\n\nThis is where I have seen the most confident nonsense. On a React Native VPN app, the real behaviour lives in Android's `VpnService` and iOS Network Extensions, with entitlements, separate processes and background rules that differ per OS and per version (I went into that split in [A React Native VPN Is Two Programs, and Only One of Them Is JavaScript](https://dev.to/nabeelbaghoor/a-react-native-vpn-is-two-programs-and-only-one-of-them-is-javascript-475j)).\n\nAI tools will happily produce a method that does not exist, a permission flow that was deprecated two versions ago, or a lifecycle assumption that only fails on a real device after the app has been backgrounded for ten minutes. None of those fail at compile time in an obvious way. Some of them do not fail until a user files a ticket.\n\nI still use the tool here, but as a fast reader of documentation that I then check myself, never as the author of code I accept on trust. The same went for the cross-platform C# networking core at Geonode: the shared logic was fine to generate against, the per-platform edges were not.\n\nUnity serialises scenes and prefabs as YAML, which looks exactly like text an AI can edit. In practice those files are full of object IDs and references the editor manages, and a hand edit that is syntactically fine can quietly break a reference you only notice at runtime.\n\nSo the line I draw is about ownership of the file format, not about any one model. I let the tool write C# scripts, editor tooling and build scripts. I do not let it edit what the Unity editor owns. If a human should not be hand-editing a file, neither should a model.\n\nIn AR, \"does it work\" often means \"does it line up with the real world when you stand there holding the phone\". No coding tool can see that. The coordinate maths can be generated, but whether an anchor drifts half a metre on a real device is a field test, not a code review.\n\nSome things I do not delegate, however good the tools get:\n\nThe honest version is that AI tools make me faster on the parts of a project that were never the hard part: wiring, boilerplate, config, test scaffolding, reading unfamiliar code. That time goes back into what actually decides whether the product works, which is architecture, edge cases, platform-specific behaviour and conversations about what the client really needs.\n\nWhat they do not replace is the person who knows why a Unity reference broke, why an iOS extension got killed in the background, or why a voice agent should transfer a call instead of answering it. If anything they raise the value of that knowledge, because they produce a lot of plausible output and someone has to know which parts are wrong.\n\nThe rule underneath all of it is short: the tool writes, I own.\n\nI am curious where other people have drawn their line. If you work on a stack that is not mostly web, which layer did you stop trusting first?\n\n*Originally published on [nabeelbaghoor.com](https://nabeelbaghoor.com/blog/claude-code-on-client-projects/).*", "url": "https://wpnews.pro/news/ai-coding-tools-are-great-until-the-only-test-is-a-device-in-someone-s-hand", "canonical_source": "https://dev.to/nabeelbaghoor/ai-coding-tools-are-great-until-the-only-test-is-a-device-in-someones-hand-3apk", "published_at": "2026-10-08 21:35:03+00:00", "updated_at": "2026-10-08 21:48:41.339919+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "ai-agents", "large-language-models"], "entities": ["Claude Code", "Cursor", "Fortell AI", "RAQTS", "Bettershop", "Tested Media", "Geonode", "ARCortex"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-coding-tools-are-great-until-the-only-test-is-a-device-in-someone-s-hand", "markdown": "https://wpnews.pro/news/ai-coding-tools-are-great-until-the-only-test-is-a-device-in-someone-s-hand.md", "text": "https://wpnews.pro/news/ai-coding-tools-are-great-until-the-only-test-is-a-device-in-someone-s-hand.txt", "jsonld": "https://wpnews.pro/news/ai-coding-tools-are-great-until-the-only-test-is-a-device-in-someone-s-hand.jsonld"}}