cd /news/ai-tools/burstiness-and-n-grams-a-40-line-pyt… · home › topics › ai-tools › article
[ARTICLE · art-142412] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Burstiness and n-grams: a 40-line Python AI-text detector, and where it breaks

A developer built a 40-line Python detector that scores text as AI-generated using burstiness (coefficient of variation in sentence length) and repeated n-grams, and reported that it fails in three ways: short texts under ~40 words swing wildly, uniform genres like legal summaries and API docs get flagged despite human authorship, and simple rewrites such as splitting a sentence or swapping "utilize" for "use" zero out the n-gram signal. The author concluded the detector works only as a triage smoke test, not a judge of writing quality, and that fixed rewrite moves — varying sentence openings, cutting redundant "that", replacing noun-stacks with verbs, and breaking rhythm every third sentence — did more to change how output reads.

by read3 min views2 publishedSep 30, 2026

Every writer who pastes drafts into an AI humanizer eventually asks the same thing: how do I know the text actually sounds human before I send it? I built a small detector to answer that, and the interesting part was not the code — it was watching where it confidently fooled itself.

Here is the whole thing, runnable as-is.

import re
from collections import Counter

def sentences(text):
    return [s.strip() for s in re.split(r'(?<=[.!?])\s+', text) if s.strip()]

def burstiness(text):
    lens = [len(s.split()) for s in sentences(text)]
    if len(lens) < 3:
        return None  # too short to mean anything
    mean = sum(lens) / len(lens)
    var = sum((x - mean) ** 2 for x in lens) / len(lens)
    return var ** 0.5 / mean  # coefficient of variation

def repeated_ngrams(text, n=3, threshold=2):
    words = re.findall(r"[a-z']+", text.lower())
    grams = [" ".join(words[i:i+n]) for i in range(len(words)-n+1)]
    return {g: c for g, c in Counter(grams).items() if c >= threshold}

def score(text):
    b = burstiness(text)
    rep = repeated_ngrams(text)
    if b is None:
        return None
    return round((1 - min(b, 1)) * 0.6 + min(len(rep), 10) / 10 * 0.4, 3)

if __name__ == "__main__":
    human = "I shipped the proxy on a Tuesday. Two days later the stream broke. Turned out the client buffered wrong. Fixed it by Friday."
    robot = "The implementation of the proxy was completed successfully. The system was designed to be robust. The architecture ensures reliability. The solution provides scalability. The approach demonstrates effectiveness."
    print("human:", score(human))   # usually higher (varied lengths)
    print("robot:", score(robot))   # usually lower (uniform, repetitive)

Run it and you will see the robot sample scores lower almost every time. Good enough to sort a pile of drafts, right? Three things broke that assumption in production.

Pitfall 1: short texts are noise. burstiness returns None under 3 sentences, and even at 3–4 sentences the coefficient of variation swings wildly. A two-sentence human reply can look more "AI" than a ten-sentence marketing page. I had to treat anything under ~40 words as "unknown" instead of scoring it.

Pitfall 2: genre beats author. Legal summaries, release notes, and API docs are supposed to be uniform. My detector flagged real human technical writing as AI because the domain is naturally low-burstiness. Burstiness alone is a genre signal, not an authorship signal.

Pitfall 3: the detector is the easiest thing to game. Swapping two sentences, splitting one long one into two, or changing "utilize" to "use" drops the repeated-ngram count to zero. That is exactly what a cheap humanizer does — it moves the score without moving the prose.

So detection helped me triage, but it never told me the text was good. The part that actually changed how the output reads was a fixed set of rewrite moves: vary sentence openings, cut the second "that", replace noun-stacks with verbs, and break the rhythm every third sentence. We baked those into a tool we use daily — ShipCopy — but the patterns work fine by hand once you have seen them fail a detector a few times.

The takeaway: a 40-line detector is a fine smoke test. It is not a judge of quality, and anyone who promises the text is human is selling the score, not the sentence.

── more in #ai-tools 4 stories · sorted by recency
── more on @shipcopy 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/burstiness-and-n-gra…] indexed:0 read:3min 2026-09-30 · —