{"slug": "burstiness-and-n-grams-a-40-line-python-ai-text-detector-and-where-it-breaks", "title": "Burstiness and n-grams: a 40-line Python AI-text detector, and where it breaks", "summary": "A developer built a 40-line Python detector that scores text as AI-generated using burstiness (coefficient of variation in sentence length) and repeated n-grams, and reported that it fails in three ways: short texts under ~40 words swing wildly, uniform genres like legal summaries and API docs get flagged despite human authorship, and simple rewrites such as splitting a sentence or swapping \"utilize\" for \"use\" zero out the n-gram signal. The author concluded the detector works only as a triage smoke test, not a judge of writing quality, and that fixed rewrite moves — varying sentence openings, cutting redundant \"that\", replacing noun-stacks with verbs, and breaking rhythm every third sentence — did more to change how output reads.", "body_md": "Every writer who pastes drafts into an AI humanizer eventually asks the same thing: how do I know the text actually sounds human before I send it? I built a small detector to answer that, and the interesting part was not the code — it was watching where it confidently fooled itself.\n\nHere is the whole thing, runnable as-is.\n\n``` python\nimport re\nfrom collections import Counter\n\ndef sentences(text):\n    # split on . ! ? followed by whitespace; keeps the delimiter\n    return [s.strip() for s in re.split(r'(?<=[.!?])\\s+', text) if s.strip()]\n\ndef burstiness(text):\n    lens = [len(s.split()) for s in sentences(text)]\n    if len(lens) < 3:\n        return None  # too short to mean anything\n    mean = sum(lens) / len(lens)\n    var = sum((x - mean) ** 2 for x in lens) / len(lens)\n    return var ** 0.5 / mean  # coefficient of variation\n\ndef repeated_ngrams(text, n=3, threshold=2):\n    words = re.findall(r\"[a-z']+\", text.lower())\n    grams = [\" \".join(words[i:i+n]) for i in range(len(words)-n+1)]\n    return {g: c for g, c in Counter(grams).items() if c >= threshold}\n\ndef score(text):\n    b = burstiness(text)\n    rep = repeated_ngrams(text)\n    # low burstiness + repeated phrases => suspicious\n    if b is None:\n        return None\n    return round((1 - min(b, 1)) * 0.6 + min(len(rep), 10) / 10 * 0.4, 3)\n\nif __name__ == \"__main__\":\n    human = \"I shipped the proxy on a Tuesday. Two days later the stream broke. Turned out the client buffered wrong. Fixed it by Friday.\"\n    robot = \"The implementation of the proxy was completed successfully. The system was designed to be robust. The architecture ensures reliability. The solution provides scalability. The approach demonstrates effectiveness.\"\n    print(\"human:\", score(human))   # usually higher (varied lengths)\n    print(\"robot:\", score(robot))   # usually lower (uniform, repetitive)\n```\n\nRun it and you will see the robot sample scores lower almost every time. Good enough to sort a pile of drafts, right? Three things broke that assumption in production.\n\n**Pitfall 1: short texts are noise.** `burstiness` returns `None` under 3 sentences, and even at 3–4 sentences the coefficient of variation swings wildly. A two-sentence human reply can look more \"AI\" than a ten-sentence marketing page. I had to treat anything under ~40 words as \"unknown\" instead of scoring it.\n\n**Pitfall 2: genre beats author.** Legal summaries, release notes, and API docs are supposed to be uniform. My detector flagged real human technical writing as AI because the domain is naturally low-burstiness. Burstiness alone is a genre signal, not an authorship signal.\n\n**Pitfall 3: the detector is the easiest thing to game.** Swapping two sentences, splitting one long one into two, or changing \"utilize\" to \"use\" drops the repeated-ngram count to zero. That is exactly what a cheap humanizer does — it moves the score without moving the prose.\n\nSo detection helped me triage, but it never told me the text was *good*. The part that actually changed how the output reads was a fixed set of rewrite moves: vary sentence openings, cut the second \"that\", replace noun-stacks with verbs, and break the rhythm every third sentence. We baked those into a tool we use daily — [ShipCopy](https://launchcraft.io/shipcopy) — but the patterns work fine by hand once you have seen them fail a detector a few times.\n\nThe takeaway: a 40-line detector is a fine smoke test. It is not a judge of quality, and anyone who promises the text is human is selling the score, not the sentence.", "url": "https://wpnews.pro/news/burstiness-and-n-grams-a-40-line-python-ai-text-detector-and-where-it-breaks", "canonical_source": "https://dev.to/keheai_harvey/burstiness-and-n-grams-a-40-line-python-ai-text-detector-and-where-it-breaks-25gl", "published_at": "2026-09-30 10:07:22+00:00", "updated_at": "2026-09-30 10:17:33.024651+00:00", "lang": "en", "topics": ["ai-tools", "natural-language-processing", "generative-ai"], "entities": ["ShipCopy", "LaunchCraft"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/burstiness-and-n-grams-a-40-line-python-ai-text-detector-and-where-it-breaks", "markdown": "https://wpnews.pro/news/burstiness-and-n-grams-a-40-line-python-ai-text-detector-and-where-it-breaks.md", "text": "https://wpnews.pro/news/burstiness-and-n-grams-a-40-line-python-ai-text-detector-and-where-it-breaks.txt", "jsonld": "https://wpnews.pro/news/burstiness-and-n-grams-a-40-line-python-ai-text-detector-and-where-it-breaks.jsonld"}}