cd /news/artificial-intelligence/ai-operations-is-a-mirage · home › topics › artificial-intelligence › article
[ARTICLE · art-119716] src=chrisarmstrong.dev ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

AI operations is a mirage

Chris Armstrong, writing on his blog, argues that AI operations is a mirage, asserting that while AI coding shows promise, applying the same logic to production operations is fraught with risks and inefficiencies. He warns that AI agents in production environments can hallucinate, make unfounded assumptions, and cause costly errors, and that current LLM-based troubleshooting often wastes time by jumping to conclusions. Armstrong advises caution, extensive harness engineering, and manual verification over autonomous agent actions.

read4 min views40 publishedAug 24, 2026
AI operations is a mirage
Image: Chrisarmstrong (auto-discovered)
  • Published on

  • Authors

  • Name

  • Chris Armstrong

I'm fairly bullish on AI coding, but my god are we so far from any kind of AI operations.

AIs are remarkably good at putting together code, being able to take well-defined plan artefacts and turn them into working systems. They still demand a high degree of architectural guidance to stop a system from degenerating into structural slop, but much of the low level work — writing tests, carrying out code reviews, responding to type errors and linting feedback — can be largely automated, albeit with a high cognitive tax on the human reviewing it.

However, projecting the same logic to operations leaves us in a somewhat different place. AI-based coding has guardrails: human code reviewers, validation in non-production environments, a cheap undo in the form of source control. Every decision has a second chance that allows you to back out, a luxury you do not have with a production system that isn't designed in the same way. Closing that gap is going to cost a lot more than you expect.

The same failures that apply to agentic development — hallucinations, making assumptions about the operational environment without checking, inconsistencies in scripts, overstepping implied boundaries, eagerness — are ten times worse when your agent can make changes in your production environment.

And that's with an agent that can't touch anything. Using an agent to even generate instructions for me has me triple checking everything, getting it to make corrections, clarifications, refinements — a process just as slow as doing it myself, except with less reading of the documentation.

My current role also has me using LLMs to do a fair bit of troubleshooting and problem diagnosis, which is much safer. The biggest issue there is timewasting: agents jump to conclusions, don't test hypotheses, over-estimate impact, and leave you with a todo list where most of it is irrelevant.

A recurring issue I see is metrics patterns that indicate an overloaded database: high CPU, instances running out of memory. These are SQL databases, where the workload can be quite unpredictable because they allow flexible data storage and access patterns. 1 Depending on the complexity of the query, size of the table, decisions of your access planner, indexes defined, whether background processes like VACUUM

are running, one query can easily overload an entire database — and this behaviour can vary depending on the other load in the database too.

2Because of this variation, an LLM investigating the issue could get sidetracked by something else it spots going on at the same time, it could fixate on something immaterial and more of a red herring, or it could simply make assumptions about the behaviour of a query without testing its impact or even running a query to see the access plan for it.

These failures cost enough to question whether using an LLM is a time-saving device at all, especially for those who know better than their agent most of the time. The problem is worse when you don't know the environment or the application you're investigating — especially the history of incidents and idiosyncrasies, the things that are often not written down. And even if you could feed them to an LLM, they're another opportunity to poison the context and pursue a dead end.

So: proceed with caution. Do a shitton of harness engineering — sandboxing, limited roles, context control, skills focussed narrowly on specific tasks. Prefer simple instructions you can run manually and verify over scripts. Prefer markdown outputs over interactive artefacts: agents work better with them, and you're not paying for a nice browser report with a dumber agent making more mistakes.

Unless your system is backed up regularly to an account the LLM can't touch, you have working and well-tested recovery patterns, your release process is smooth and frequent, tests run comprehensively and without flakiness, and you have the operational discipline to use feature flags and other blast-radius reducing change management practices, you should tread very carefully before using LLMs for operational concerns. Treat the output with a high degree of skepticism, and think twice before giving them administrative access.

Companies — whether they know it or not — require much more DevOps and platform engineering than they did in the past, because they're trying to move faster with architectures and processes that were insufficient before AI coding and are now downright dangerous.

The need for operational expertise hasn't gone away. Quite horribly, I think it's become much more in demand — at exactly the time when more people can generate poor results while playing expert.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @chris armstrong 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-operations-is-a-m…] indexed:0 read:4min 2026-08-24 · —