cd /news/ai-tools/i-had-ai-grade-my-12-vibe-coded-blog… · home › topics › ai-tools › article
[ARTICLE · art-145732] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

I Had AI Grade My 12 Vibe-Coded Blog Posts: 15.1 out of 25

A developer audited 12 AI-assisted blog posts using two fresh-context subagents, scoring an average of 15.1 out of 25 across uniqueness, specificity, accuracy, structure, and authenticity. The most common defect was a formulaic fake-anecdote sentence shape repeated across four posts, which the developer caught with a regex and moved into an 11-check quality gate script (cc_quality_gate.py), raising re-audit scores to 19 and then 21. The developer notes that contradictions between posts and stale model names remain undetectable by pattern matching and require human review.

by read5 min views1 publishedOct 5, 2026

On 2026-08-16 I had two fresh-context subagents audit 12 posts from my own AI-assisted blog, six posts each, cross-checked. The average score was 15.1 out of 25. The biggest finding was not a factual error. It was a template: four posts used almost the same fake-anecdote sentence shape ("I once ran into X, and I realized Y"). Because the defect is a repeated shape, a regex can catch it, so I moved the audit items into a quality gate script and fixed the flagged sentences with a replacement list. Re-audits scored 19 and then 21.

This post covers the audit design, the gate, the patch format, and what the gate cannot catch. The full write-up, in Korean, is in the original post on my blog.

Five axes, five points each, 25 total: uniqueness, specificity, accuracy and timeliness, structure, and authenticity. Each auditor got three instructions:

The quote requirement matters. Without it, an LLM auditor drifts into vague praise or vague criticism. With it, every finding points at a sentence you can open and check.

Using fresh-context subagents also matters. The model that wrote a post tends to approve of it. Two auditors that never saw the drafting session gave a much harsher read.

Four defect types came up repeatedly:

Defect Example Machine-detectable
Unsourced precise number "62%" reused for three different claims Partly
Formulaic fake anecdote Same sentence shape in 4 posts Yes
Contradiction inside the site "about 20,000 KRW a month" vs "about 29,000 KRW" No
Stale model names Old model described as the latest No

Two of the four cannot be caught by pattern matching. A contradiction between two posts needs both posts in view, and a stale model name needs to know what is current. Those stay with the human or the next audit.

For the defects that repeat in form, I wrote scripts/publish/cc_quality_gate.py with 11 checks: five metadata fields filled, exactly one H1 under 55 characters, a three-line summary block, a conclusion table right after it, at least four H2 sections with 60% or more phrased as questions, exactly five FAQ items, at least two tables, at least two in-sentence internal links, zero forbidden elements, real-use elements (two or more code blocks with language tags, steps, a limits paragraph), and a cap on number-bearing sentences without an evidence marker.

The two detectors that did the most work are plain regexes:

import re

ORG = re.compile(
    r"Acme Research|Example Institute|Sample University|Foo Consulting|"
    r"Annual Trend Index|Global AI Index"
)
ANECDOTE = re.compile(
    r"in my experience|I tried it and|I once|I was surprised|I realized"
)

def scan(text: str) -> list[str]:
    hits = []
    for name, rx in (("org-stat", ORG), ("fake-anecdote", ANECDOTE)):
        for m in rx.finditer(text):
            hits.append(f"{name}: {m.group(0)!r} at {m.start()}")
    return hits

My real patterns are Korean; the ones above are an English illustration of the same idea. Findings are split into errors, which stop publishing (missing metadata, wrong H1 count, wrong FAQ count, a detected institutional statistic, anecdote or emoji), and warnings, which do not (long title, few H2s, low question ratio, too few tables).

Patches are applied only through a replacement list, and each old string must occur in the source exactly once. Zero matches or two or more matches is recorded as a failure and skipped. An empty new deletes the sentence.

{"id": 126, "slug": "example-slug",
 "patches": [
   {"old": "An Acme Research report says 40% improved", "new": "", "why": "unsourced institutional stat"},
   {"old": "I tried it and was surprised", "new": "The measured result is below", "why": "invented anecdote"}
 ]}
php
def apply(text: str, patches: list[dict]) -> tuple[str, list[str]]:
    failed = []
    for p in patches:
        if text.count(p["old"]) != 1:
            failed.append(p["why"])
            continue
        text = text.replace(p["old"], p["new"])
    return text, failed

The exactly-once rule means a patch can never silently rewrite the wrong paragraph. A skipped patch is visible in the failure list, and a human looks at it.

The first re-audit scored 19 out of 25, under my 20 point bar, so I fixed the flagged items first. The second re-audit scored 21 out of 25. The operating rules after that: the target is 20 or higher, anything below blocks publishing until the flagged items are fixed, and if three months after publishing the score keeps falling under 18, expansion stops and I re-diagnose the cause.

The rubric for the next audit anchors each axis. Uniqueness is 1 for content found anywhere, 3 for generic advice with some own examples, 5 for three or more file paths or measured values. Specificity is 5 only when a reader can copy and run the command or code. Authenticity is 5 when misdiagnoses and failures are written up with their causes.

Treat the numbers as relative. They come from the same harness measured again, not from an absolute scale.

The gate only catches defects with a repeated form. "According to one study" has no institution name, so the organization regex misses it. A new anecdote phrasing that is not on the list passes. Contradictions across posts and outdated model names are invisible to it.

The rule I took from this is short: write only what you measured, what actually broke on you, or what exists as a file. Also save each post's total score to a file; a session transcript alone is weak evidence of what was scored.

If you run AI-assisted writing at any volume, an audit pass with fresh contexts costs little and finds patterns your own review will skip. The checklist, the rubric, and the Korean regexes are in the full post.

Free tools I keep on my own site: 34 calculators and how-to guides - loans, severance pay, take-home salary, savings vs deposits. New ones go out in one email, no ads.

── more in #ai-tools 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-had-ai-grade-my-12…] indexed:0 read:5min 2026-10-05 · —