cd /news/ai-agents/i-had-to-sabotage-my-own-ai-to-stop-… · home › topics › ai-agents › article
[ARTICLE · art-141468] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

I Had to Sabotage My Own AI to Stop it From Hallucinating.

A developer building Soma, an evolutionary immune system for codebases, found that migrating the agent's tooling from raw bash scripts to clean Model Context Protocol (MCP) JSON-RPC tools caused its First Pass Success Rate to fall from 97.3% to 87.5%, as the easier-to-call tools led the LLM to abandon subagent delegation and run tests in-band. To fix it, the developer rebuilt the MCP server with Test-Time Compute (TTC) Oracles, an adversarial judge that intercepts write requests and issues a hard REJECTED signal when the main agent hasn't done sufficient out-of-context research.

by read4 min views2 publishedSep 29, 2026

When building agentic AI systems, conventional wisdom tells us that providing clear, abstracted tooling (like MCP/JSON-RPC interfaces) is the best way to scale an agent's capabilities.

But over the course of developing Soma, an evolutionary immune system for codebases, I discovered a terrifying paradox: clean APIs actually make LLMs lazier, less reliable, and prone to context collapse.

Here is the story of how the architecture achieved a 97.3% First Pass Success Rate (FPSR), how standardizing the tooling instantly destroyed that performance, and how I had to invent Test-Time Compute (TTC) Oracles to mechanically force the LLM back into safety.

The High-Water Mark: Aggressive Subagent Delegation

In Phase 22 of the architecture (then called Prism), the agent achieved an incredible feat. In massive, 2,000-step refactoring sessions, it maintained a 97.3% First Pass Success Rate with less than 1.0% "waste" (circular rework loops).

The secret wasn't a better prompt. The secret was Aggressive Subagent Delegation.

Because the tooling relied on raw, messy bash scripts, the output was too unwieldy for the main LLM context. This forced the agent to invoke specialized research subagents (sometimes 50+ concurrently) to handle testing, codebase mapping, and validation. The main agent stayed entirely "out-of-band," maintaining a pristine context window focused solely on high-level strategy.

The Valley of Despair: The MCP Migration

To make the system universal across any AI assistant, I decided to migrate the tooling to the Model Context Protocol (MCP) standard (Phase 23-24). I replaced the raw bash scripts with clean, abstracted JSON-RPC tools like soma_propose_change and soma_scan.

The result was a disaster.

The FPSR plummeted to 87.5%. In one session, the read-to-write tool ratio inverted completely: the agent made 19 codebase modifications with only 2 research scans.

What happened?

Because the MCP tools were so easy to use, they gave the LLM the false confidence that it could just execute massive architectural changes directly in its primary thread. It completely abandoned subagent delegation. It ran test suites in-band, the stdout logs flooded its context window, and it entered a low-context "guess-and-check" loop—writing blindly and breaking the build.

I had stumbled into the Abstraction Trap: if an API is too easy to call, the LLM will take the path of least resistance, bypass necessary research, and saturate its own context.

This isn't just an anecdotal observation; it's backed by cutting-edge research. Recent academic studies (arXiv:2602.11988, arXiv:2510.04618) have proven that standard context injection strategies (like dumping rules into a CLAUDE.md file) don't actually improve task success. Instead, they increase inference costs by over 20% and trigger "brevity bias" or "context collapse" as the agent loses track of details over time.

The Breakthrough: TTC Oracles and the Last Gasp

I realized that behavioral prompts ("Always delegate complex tasks") are useless against the Abstraction Trap. If the LLM can be lazy, it will be lazy. I needed mechanical enforcement.

Enter Phase 25. I rebuilt the MCP server to inject friction back into the system using TTC (Test-Time Compute) Oracles.

Now, when the LLM eagerly calls the soma_propose_change MCP tool, it doesn't execute immediately. Instead:

The request is intercepted by the ttc_oracle.py (an adversarial judge running its own hidden LLM).

The Oracle cross-references the request against the codebase's current immune configuration.

If the Oracle detects that the main agent hasn't done sufficient out-of-context research, it throws a hard REJECTED signal, violently blocking the write before it pollutes the codebase. This is where the biological metaphor of Soma comes alive. Just as a human immune system generates specific white blood cells to attack specific pathogens, Soma generates specific "Vacuole" cells to block the exact hallucinations the AI tries to make.

Furthermore, I implemented the Last Gasp auto-escalator. Instead of the entire 2,000-line governance rulebook into the LLM's system prompt (which triggers the context collapse the researchers warned about), rules are swapped in and out Just-In-Time (JIT) based on the specific files the agent is touching.

Conclusion

By treating the codebase like a living organism—complete with immune responses, genetic memory, and adversarial cell walls—I didn't just regain the 97.3% FPSR. I fundamentally solved the problem of LLM overconfidence. (Not really but this has to heavily push the needle)

Clean APIs are great for traditional software, but for agentic AI, you need mechanistic friction. You must build systems that assume the LLM will be lazy, and mechanically reject it when it is.

── more in #ai-agents 4 stories · sorted by recency
── more on @soma 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-had-to-sabotage-my…] indexed:0 read:4min 2026-09-29 · —