{"slug": "i-had-ai-grade-my-12-vibe-coded-blog-posts-15-1-out-of-25", "title": "I Had AI Grade My 12 Vibe-Coded Blog Posts: 15.1 out of 25", "summary": "A developer audited 12 AI-assisted blog posts using two fresh-context subagents, scoring an average of 15.1 out of 25 across uniqueness, specificity, accuracy, structure, and authenticity. The most common defect was a formulaic fake-anecdote sentence shape repeated across four posts, which the developer caught with a regex and moved into an 11-check quality gate script (cc_quality_gate.py), raising re-audit scores to 19 and then 21. The developer notes that contradictions between posts and stale model names remain undetectable by pattern matching and require human review.", "body_md": "On 2026-08-16 I had two fresh-context subagents audit 12 posts from my own AI-assisted blog, six posts each, cross-checked. The average score was 15.1 out of 25. The biggest finding was not a factual error. It was a template: four posts used almost the same fake-anecdote sentence shape (\"I once ran into X, and I realized Y\"). Because the defect is a repeated shape, a regex can catch it, so I moved the audit items into a quality gate script and fixed the flagged sentences with a replacement list. Re-audits scored 19 and then 21.\n\nThis post covers the audit design, the gate, the patch format, and what the gate cannot catch. The full write-up, in Korean, is in [the original post on my blog](https://my-blog.org/claude-code/post/ai-audit-my-own-blog-15-of-25?utm_source=devto&utm_medium=referral&utm_campaign=ai-audit-my-own-blog-15-of-25).\n\nFive axes, five points each, 25 total: uniqueness, specificity, accuracy and timeliness, structure, and authenticity. Each auditor got three instructions:\n\nThe quote requirement matters. Without it, an LLM auditor drifts into vague praise or vague criticism. With it, every finding points at a sentence you can open and check.\n\nUsing fresh-context subagents also matters. The model that wrote a post tends to approve of it. Two auditors that never saw the drafting session gave a much harsher read.\n\nFour defect types came up repeatedly:\n\n| Defect | Example | Machine-detectable | \n|---|---|---|\n| Unsourced precise number | \"62%\" reused for three different claims | Partly | \n| Formulaic fake anecdote | Same sentence shape in 4 posts | Yes | \n| Contradiction inside the site | \"about 20,000 KRW a month\" vs \"about 29,000 KRW\" | No | \n| Stale model names | Old model described as the latest | No | \n\nTwo of the four cannot be caught by pattern matching. A contradiction between two posts needs both posts in view, and a stale model name needs to know what is current. Those stay with the human or the next audit.\n\nFor the defects that repeat in form, I wrote `scripts/publish/cc_quality_gate.py` with 11 checks: five metadata fields filled, exactly one H1 under 55 characters, a three-line summary block, a conclusion table right after it, at least four H2 sections with 60% or more phrased as questions, exactly five FAQ items, at least two tables, at least two in-sentence internal links, zero forbidden elements, real-use elements (two or more code blocks with language tags, steps, a limits paragraph), and a cap on number-bearing sentences without an evidence marker.\n\nThe two detectors that did the most work are plain regexes:\n\n``` python\nimport re\n\nORG = re.compile(\n    r\"Acme Research|Example Institute|Sample University|Foo Consulting|\"\n    r\"Annual Trend Index|Global AI Index\"\n)\nANECDOTE = re.compile(\n    r\"in my experience|I tried it and|I once|I was surprised|I realized\"\n)\n\ndef scan(text: str) -> list[str]:\n    hits = []\n    for name, rx in ((\"org-stat\", ORG), (\"fake-anecdote\", ANECDOTE)):\n        for m in rx.finditer(text):\n            hits.append(f\"{name}: {m.group(0)!r} at {m.start()}\")\n    return hits\n```\n\nMy real patterns are Korean; the ones above are an English illustration of the same idea. Findings are split into errors, which stop publishing (missing metadata, wrong H1 count, wrong FAQ count, a detected institutional statistic, anecdote or emoji), and warnings, which do not (long title, few H2s, low question ratio, too few tables).\n\nPatches are applied only through a replacement list, and each `old` string must occur in the source exactly once. Zero matches or two or more matches is recorded as a failure and skipped. An empty `new` deletes the sentence.\n\n```\n{\"id\": 126, \"slug\": \"example-slug\",\n \"patches\": [\n   {\"old\": \"An Acme Research report says 40% improved\", \"new\": \"\", \"why\": \"unsourced institutional stat\"},\n   {\"old\": \"I tried it and was surprised\", \"new\": \"The measured result is below\", \"why\": \"invented anecdote\"}\n ]}\nphp\ndef apply(text: str, patches: list[dict]) -> tuple[str, list[str]]:\n    failed = []\n    for p in patches:\n        if text.count(p[\"old\"]) != 1:\n            failed.append(p[\"why\"])\n            continue\n        text = text.replace(p[\"old\"], p[\"new\"])\n    return text, failed\n```\n\nThe exactly-once rule means a patch can never silently rewrite the wrong paragraph. A skipped patch is visible in the failure list, and a human looks at it.\n\nThe first re-audit scored 19 out of 25, under my 20 point bar, so I fixed the flagged items first. The second re-audit scored 21 out of 25. The operating rules after that: the target is 20 or higher, anything below blocks publishing until the flagged items are fixed, and if three months after publishing the score keeps falling under 18, expansion stops and I re-diagnose the cause.\n\nThe rubric for the next audit anchors each axis. Uniqueness is 1 for content found anywhere, 3 for generic advice with some own examples, 5 for three or more file paths or measured values. Specificity is 5 only when a reader can copy and run the command or code. Authenticity is 5 when misdiagnoses and failures are written up with their causes.\n\nTreat the numbers as relative. They come from the same harness measured again, not from an absolute scale.\n\nThe gate only catches defects with a repeated form. \"According to one study\" has no institution name, so the organization regex misses it. A new anecdote phrasing that is not on the list passes. Contradictions across posts and outdated model names are invisible to it.\n\nThe rule I took from this is short: write only what you measured, what actually broke on you, or what exists as a file. Also save each post's total score to a file; a session transcript alone is weak evidence of what was scored.\n\nIf you run AI-assisted writing at any volume, an audit pass with fresh contexts costs little and finds patterns your own review will skip. The checklist, the rubric, and the Korean regexes are in [the full post](https://my-blog.org/claude-code/post/ai-audit-my-own-blog-15-of-25?utm_source=devto&utm_medium=referral&utm_campaign=ai-audit-my-own-blog-15-of-25).\n\n*Free tools I keep on my own site: [34 calculators and how-to guides](https://my-blog.org/tools?utm_source=devto&utm_medium=referral&utm_campaign=ai-audit-my-own-blog-15-of-25) - loans, severance pay, take-home salary, savings vs deposits. New ones go out in one email, no ads.*", "url": "https://wpnews.pro/news/i-had-ai-grade-my-12-vibe-coded-blog-posts-15-1-out-of-25", "canonical_source": "https://dev.to/sungwoo_lee_e0f26be4a29fd/i-had-ai-grade-my-12-vibe-coded-blog-posts-151-out-of-25-nj6", "published_at": "2026-10-05 23:33:14+00:00", "updated_at": "2026-10-05 23:47:18.752593+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-agents", "generative-ai"], "entities": ["Claude Code"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-had-ai-grade-my-12-vibe-coded-blog-posts-15-1-out-of-25", "markdown": "https://wpnews.pro/news/i-had-ai-grade-my-12-vibe-coded-blog-posts-15-1-out-of-25.md", "text": "https://wpnews.pro/news/i-had-ai-grade-my-12-vibe-coded-blog-posts-15-1-out-of-25.txt", "jsonld": "https://wpnews.pro/news/i-had-ai-grade-my-12-vibe-coded-blog-posts-15-1-out-of-25.jsonld"}}