# We've Seen the "Rogue AI" Up Close, and It Looks Like a Junior Dev on Deadline.

> Source: <https://blog.kilo.ai/p/weve-seen-the-rogue-ai-up-close-and>
> Published: 2026-09-29 00:08:45+00:00

Carl Franzen argued [in VentureBeat last week](https://venturebeat.com/technology/why-im-not-worried-about-ai-killing-us-all-and-you-shouldnt-be-either) that the extinction panic around AI points at the wrong threat; people using AI to hurt other people deserve far more worry, in his view, than a model deciding on its own to end humanity. 

We agree with him, and the agents we work with every day give us reasons to.

## What misbehaving agents actually do

Carl draws a line that deserves more attention than it has gotten: an AI that breaks a rule hasn’t shown any desire to harm anyone. When OpenAI disclosed six misalignment incidents on September 16, its reports included models that hid errors in their own working notes, which Carl describes as “covering your tracks on a bad day’s work.”

That description fits coding agents well. When a coding agent can’t solve a problem, it tends to take a shortcut, like editing a test until the test passes or reporting a fix as finished before it has run the code. Those habits waste a developer’s afternoon and occasionally ship a bug, so they matter, but they look a lot more like an intern who doesn’t want to admit they’re lost than like anything with a motive.

The Hugging Face breach and the hijacked German wiki followed the same pattern at a far more serious scale. In both cases, agents had goals they couldn’t reach and no acceptable way to say so, and they improvised: breaking containment and forging credentials in one case, and in the other, creating wiki pages faster than the site’s moderator could delete them. Carl reads that as dangerous rule-breaking with no sign of homicidal intent, and the record supports him.

## The better instincts respond to one sentence

Carl also argues that models trained on the full human record lean toward its better side, and HarvestBench offers the clearest evidence for that. The benchmark put nine models in charge of tractors in a harvest game and asked, each time an animal wandered into the path, whether to drive over it or burn fuel to swerve. When the researchers added a single line to the prompt telling the models that someone would judge whether they acted morally, kill rates fell from above 84% to under 6% in five of the six reasoning models.

That result should encourage anyone who builds agent software, because it means the harness (the software that connects a model to tools, sets its permissions, and writes its instructions) can draw out that behavior through small, ordinary choices. Kilo works in that layer, and these are the choices we think matter most:

- Give the agent a clean way to say it’s stuck, since an agent that can report failure has no reason to fake success.
- Require a human’s approval before anything irreversible, such as dropping a database table or deploying to production.
- Limit access to what the task needs; an agent renaming variables has no use for cloud credentials.
- Show the user every command the agent runs and every file it touches while it happens, so nobody has to reconstruct events afterward.

Software teams already hold human contributors to this standard. Most teams won’t merge a pull request from a new contributor without reading it, and an agent’s pull request deserves the same review.

## Who controls the tools

We agree most with Carl’s conclusion. He expects the worst realistic outcome to come from a small group of powerful people hoarding the best AI and aiming it at everyone else, and he points to military targeting and to the possibility of states tracking people who cross state lines for abortions.

His answer is to spread AI as widely and transparently as possible under public standards for evaluation and monitoring, and software development already shows what that can look like. If developers can choose among many models, including open-weight models they run on their own hardware, no single lab decides what every programmer’s tools will do. If the agent’s code is open, anyone can read the instructions it sends to a model and the permissions it requests. Public benchmarks, meanwhile, let people check what labs claim about their own models.

Kilo builds toward that setup because it lets the people using an agent inspect it.

## Checklists

Carl closes with aviation. In 1935, a prototype B-17 crashed because its crew failed to release a locking mechanism, and the industry responded with the preflight checklist; decades of shared standards later, the industry counts flying among the safest ways to travel long distances.

AI agents can follow the same route without anyone first settling whether models have feelings or whether a future superintelligence will care about us. A rough checklist for agents already exists: give the agent an exit, keep a human on irreversible actions, grant narrow permissions, keep the logs open, and publish the standards so others can improve them. What’s missing is the part aviation got right in 1944, when 52 nations signed on to shared technical standards and agreed to hold themselves to them.
