# Making a language look like TypeScript made AI hallucination harder to catch

> Source: <https://dev.to/lolocoding/making-a-language-look-like-typescript-made-ai-hallucination-harder-to-catch-3llg>
> Published: 2026-10-06 09:11:35+00:00

The original developer-first playbook was world-class docs, beginner-friendly SDKs, and advocates at every conference. Stripe and Twilio wrote it around 2012, and the whole industry copied it, including me, for the majority of my DevRel career.

Every piece of that playbook assumed a human being would read what we wrote.

That assumption is what broke, and you have probably seen the conclusion people have been drawing from it: wind down the Developer Relations team, because AI-native tooling lets more people build without one.

I understand the instinct, I really do. Getting to a first experiment has never been cheaper, and fewer developers need someone to walk them through step one. I don't see devs reading tutorials about how to use a thing anymore. Mostly they just point their agent at the docs and ask for what they actually want.

That argument treats "AI-native tooling" as a condition of the world, like bandwidth. Something that arrived and now simply *exists*. So recently, I decided to test whether it had.

October 2026. One prompt, four attempts: write a contract for this platform that lets someone prove they are over 18 without revealing their date of birth.

Three of those attempts were models with no tooling for this platform. No plugins, no memory, no custom instructions. One of the three went and read the documentation on its own. The fourth had the tooling my team built, and a compiler on the machine.

Every file in this language opens by declaring which version of the language it was written for. Get that line wrong, and nothing underneath it matters.

| Attempt | Declared | Distance from current | 
|---|---|---|
| Gemini | `0.14` | 13 versions behind | 
| Claude, no plugins | `>= 0.16` | 11 versions behind | 
| ChatGPT, which cited its sources | `0.26` | 1 version behind | 
| Claude, with the tooling | `0.27` | current | 

The one that came closest was the one that went and read something. The one that was correct did not recall anything at all. Its first move was to distrust itself: the tooling includes an instruction telling the model its own knowledge of this language is unreliable and to check the installed compiler before writing a line. It checked, found the version, compiled clean on the first pass, and put the result through a prover.

Two of those four runs were the same underlying model (Claude). The difference was whether somebody had built the thing that made it check.

When a coding agent runs into a programming language it has never seen, it gets things right less than 20% of the time. Give it the syntax rules written out, some working examples, a way to look up what's actually available, and a compiler it can test itself against, and it climbs to 85%. Those numbers come from [Microsoft's own guidance](https://devblogs.microsoft.com/all-things-azure/ai-coding-agents-domain-specific-languages/) to teams building with niche languages.

So 65 points of accuracy sit between a platform an agent can build on, and a platform an agent makes things up about. The model is identical on both sides of that gap.

Somebody needs to teach the machine how the platform works, and then keep teaching it. That is what makes a platform AI-native.

I spent the last stretch of my career running Developer Relations for a privacy platform that had its own smart contract language. That language was built to look like TypeScript, on purpose. Web developers are the biggest pool you can hire from, and learning privacy-preserving contract patterns is hard enough without a strange and new syntax on top of it. The bet was that a familiar-looking language would get more people to try it.

For people learning it, that worked exactly the way we hoped.

Then, developers stopped writing the first draft themselves.

A model that sees TypeScript-shaped code will reach for TypeScript-shaped answers. And it did. It wrote functions with believable names, believable arguments, believable async handling, in a language that has no async and does not call them functions.😳

And this is where the design choice turns on itself. Because when you're reading syntax you don't recognize, you're careful, because you know you don't know it. But when you're reading something that looks like TypeScript you've written hundreds, if not thousands, of times, you don't check anything. Nothing looks out of the ordinary. Nothing looks wrong.

I watched this happen over and over at our hackathons. Students would lose an hour to a function that *never existed*, in a language where the thing they wanted isn't called a function at all. And they never blamed the language. They blamed themselves, their install, their Node version, their laptop. They assumed they had done something wrong and just couldn't find where.

They were reading something that looked exactly like code they already knew how to read.

Familiar syntax doesn't remove the need to catch mistakes. It moves that job off the developer and onto the tools.

The original call to model the language off of TypeScript was right for the world it was made in. Accessibility through familiarity is a good bet when a person types the first draft. Nobody made a bad decision here. The ground simply moved underneath a good one.

Which leaves something uncomfortable for anyone building a developer platform right now. The closer your syntax sits to something popular, the more convincingly a model will fabricate against it, and the less likely your developer is to catch it. Your tooling has to absorb the checking your developer is no longer doing.

Documentation written for humans wasn't going to cut it anymore.

Human docs explain, and they assume a reader who builds a mental model, notices what is missing, and goes and asks. A model does none of that. It needs the rules stated outright, the examples pinned to a version it can check against, and something that tells it when it is wrong.

So my team scoped and commissioned a set of plugins built for the AI to use, not for a person to read:

That last piece is the one people skip. The usual version of this work is a folder of examples and a hopeful README. What we needed was something that runs and reports back, and spins up a local test network so the AI looks at a real environment instead of guessing at one.

Now go back to Microsoft's list. Syntax rules. Working examples. Something to look things up in. A compiler to check against. We arrived at the same place from the other direction, by watching people fail.

By the time we got it in front of the community, it was 16 plugins, 102 skills, 17 agents, and tens of thousands of lines of reference material and example code. Humans wrote that. Humans reviewed it.

It's written against a specific compiler version. When the compiler moves on, some of it stops being true, and every plugin pointing at it goes on repeating it with complete confidence.

I don't know what any of those numbers look like today, and I'm no longer the person who can tell you. All of that material is only as good as the last time somebody checked it, and checking it is a much more specific job than it sounds.

None of the material announces that it has gone wrong. The only way to find out is to write the contract the AI would write, compile it against today's toolchain, and know enough about the language to recognize a plausible answer that happens to be false.

That takes someone who can read the code, has the compiler in front of them, and the time to do so. Most organizations hand this work to someone with just one of those. Whoever inherits it needs all of them, and needs to be asked the same question at every release:

Does it still compile?

There's a second cost, and it takes longer to show up.

A developer who hits a wall the old way files an issue, posts in the Discord, emails @support. That complaint is information. **Feedback is a gift**, and I have never been embarrassed to say so. It's how a platform team finds out where its product stops matching reality.

A developer working through an agent doesn't do any of that. The agent hands them something that looks right. It fails, so they try a different prompt, and it fails again, and they decide the platform isn't ready yet. Then they're gone. No ticket is submitted. No thread is started. Nothing will ever reach a product manager's desk.

So the same shift that brings more people in also dries up your best information about where they drop off, right at the moment you need it most.

If I were writing this from scratch, every line in it would be something you could count.

| What you build | How you know it's working | How you know it isn't | 
|---|---|---|
| Material the AI can actually use | The share of AI-assisted first attempts that compile, re-measured after every compiler release and every major model release | The rate drops, and nobody notices until a developer tells you | 
| A way to check the AI's output | Time from a developer's first prompt to a working build | Invented functions are still being found by your developers instead of by your pipeline | 
| A record of where people break | Distinct failure patterns found per quarter, and the days between finding one and a product decision | Patterns get logged, and nothing gets decided | 
| Builders who get past the first demo | How many reach production, by whatever threshold your platform actually counts | Nobody can name five without going to check | 

Again, every one of these is countable, holds up in a budget review, and can prove me wrong. If the numbers don't move, the program isn't working, and it should be cut. I'll take that deal over being judged on how excited everyone seemed.

I should clarify my stance on things. Because there are circumstances when cutting a DevRel team is absolutely the right call. Sometimes the product is two years away from needing a full team, and somebody hired one too early. Or the job got scoped as a travel schedule, and the travel stopped paying for itself.

But the work itself doesn't leave with the people. It just… moves. It lands on Support, who sees every failure, but can fix none of them. Or on Engineering, who can fix them, but has no time. Or it lands nowhere, which is a decision too, just one no one made out loud.

So before a leadership team makes that call, somebody should have to say out loud who owns the accuracy of what the AI tells developers. It can be one person, or a contractor, or a line in somebody's existing job description. It just cannot be an assumption.

Find me [@lolocoding](https://x.com/lolocoding)👩🏼💻
