cd /news/ai-agents/i-pointed-108-agents-at-my-own-growt… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-131093] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=Β· neutral

I pointed 108 agents at my own growth strategy. They refuted the two assumptions it stood on.

A developer ran a deep-research harness of 108 agents over their own content strategy, extracting 124 atomic claims and adversarially verifying 25 of them, of which 13 were confirmed and 12 refuted β€” including the two assumptions the entire plan rested on. The refuted premise was that English terminology in the author's niche had no canonical owner; a single fetch showed Wikipedia's article already defines the head term, carries the planned taxonomy, and is still actively maintained. The author argues the transferable lesson is structural: refutation must be an agent's explicit job, with the burden of proof reversed and claims surviving only by majority verifier vote.

by read5 min views1 publishedSep 16, 2026

I had a content strategy. It was coherent, it was written down, and about 60% of my planned work depended on two specific claims being true.

Then I ran a deep-research pass over it: 108 agents, 25 sources fetched, 124 claims extracted, 25 put through adversarial verification. 13 confirmed. 12 refuted.

Among the refuted twelve were the two claims the whole plan rested on. Not edge cases β€” the premise the content plan was built on, and the tactic I'd allocated the most effort to.

This post is about the machinery that produced that result, because the machinery is the transferable part. If you use an LLM to check your own thinking and it keeps agreeing with you, the problem is almost certainly your harness, not your thinking.

Ask a capable model to evaluate a plan you wrote and you get a critique β€” usually a good-sounding one, with a few "considerations" and a broadly supportive conclusion. Ask a research agent to investigate a question you've already formed a view on, and it will return sources that support the view. Not because it's sycophantic exactly, but because you framed the query, and the framing carries the answer.

I've done this to myself more than once. The output feels like diligence. It functions as confirmation.

The fix isn't a better prompt. It's making refutation somebody's job.

Four stages, and the third is the one that matters.

1. Fan out over sources, not over opinions. 25 sources fetched β€” competitor pages, the actual Wikipedia wikitext, MediaWiki revision APIs, published studies. Agents that reason about a topic from training data converge on consensus. Agents that fetch a specific artifact and report what's in it disagree productively.

2. Extract atomic claims. 124 of them. The discipline here is granularity: "our terminology positioning is sound" cannot be verified, so it isn't a claim. "English Wikipedia's Four Pillars article defines Day Master in prose" can be checked in one fetch. Big claims decompose into small checkable ones, and only the small ones enter the next stage.

3. Adversarially verify β€” with the burden of proof reversed. This is the whole thing. 25 claims went to independent verifiers whose instruction was to refute, defaulting to refuted when uncertain. Multiple verifiers per claim, majority vote, and the vote recorded with the finding:

πŸ”΄ refuted   β€” β‰₯2/3 verifiers killed it
🟒 confirmed β€” β‰₯2/3 verifiers upheld it
🟑 medium    β€” confirmed but single-source or confounded
βšͺ extracted, never verified β€” budget ran out. A lead, not a fact.

Note that a claim surviving is not the same as a claim being unexamined. 3-0 and 2-1 mean different things and I kept the split visible in the notes, because a 2-1 survival is a claim I should revisit.

4. Synthesize from the survivors only. The refuted ones don't get downgraded to "risks to monitor". They get removed, and the plan is rebuilt from the 13 that lived.

Pillar one: "English terminology in my niche has no canonical owner." This was the premise. My plan was a set of pages whose job was to define core terms.

It's false, and it took one fetch to establish. English Wikipedia's article binds the head term in its lead sentence, defines the central concept in prose, carries a 10-row table of the exact taxonomy I planned to explain β€” 4,200 words, 20 footnotes, no cleanup banners. Created 2006.

The part that actually settled it: the incumbent is still hardening. Last edited 2026-05-28, with an ~8,000-byte expansion in March, verified through the revision API rather than by eyeballing the page. I wasn't proposing to fill a vacancy. I was proposing to out-write a twenty-year-old article that gets attention every month.

Worse, published analysis of ~76.7M AI Overviews plus ~957k ChatGPT and ~953.5k Perplexity prompts shows these systems over-index on exactly encyclopedic and UGC sources. So the channel I was optimizing for is structurally biased toward the incumbent I was planning to displace.

Pillar two: "structured data will get me cited by AI systems." I had budgeted real work for JSON-LD across the site.

A matched difference-in-differences study settled it: 1,885 pages that added JSON-LD over an eight-month window, each matched to 3 control URLs on other domains at similar pre-period citation levels β€” roughly 4,000 controls β€” with a 30-day pre/post window and four statistical approaches including event-study weekly plots.

Platform Citation change after adding JSON-LD
Google AI Overviews βˆ’4.6% (small but statistically significantdecline )
Google AI Mode +2.4% β€” indistinguishable from zero
ChatGPT +2.2% β€” indistinguishable from zero

The widely-quoted correlation that AI-cited pages are ~3Γ— more likely to carry JSON-LD is confounded, and the same analysts who published it say so: schema tends to live on the kind of sites that get cited anyway.

I dropped that half of the plan. I kept the one piece that has an independent reason to exist β€” an extractable definition in the first paragraph, which helps a human reader before it helps any crawler.

Ten claims were extracted and then dropped before verification when the run hit its budget. I marked them βšͺ and moved on.

Re-reading later, the βšͺ pile turned out to hold the most decision-relevant material in the entire run β€” leads suggesting that a slot I'd assumed was open was already occupied, including by an LLM vendor's own encyclopedia. Every one of those is still unverified. I can't act on them and I can't dismiss them.

Two things follow. First: truncation is not random. Verification runs in some order, and whatever the order is, the tail is systematically different from the head β€” usually later-fetched, more specific, more surprising. Second: an unverified claim has to be visibly unverified. If those ten had been folded into the synthesis as findings, I'd have "learned" something I actually don't know. The βšͺ tag is doing more work than the πŸ”΄ tag.

The uncomfortable part: an unstructured deep-research query would have handed me a confident report agreeing with my plan, citing many of the same 25 sources. Same model, same sources, opposite conclusion. The difference was entirely in the harness.

That's what the adversarial stage buys. It's the difference between research and expensive confirmation bias β€” and it costs about one extra agent per claim.

The engine-and-model architecture those agents also poked at is documented on the method page; the product they were arguing about is auspiceoracle.com.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @wikipedia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/i-pointed-108-agents…] indexed:0 read:5min 2026-09-16 Β· β€”