# Burstiness and n-grams: a 40-line Python AI-text detector, and where it breaks

> Source: <https://dev.to/keheai_harvey/burstiness-and-n-grams-a-40-line-python-ai-text-detector-and-where-it-breaks-25gl>
> Published: 2026-09-30 10:07:22+00:00

Every writer who pastes drafts into an AI humanizer eventually asks the same thing: how do I know the text actually sounds human before I send it? I built a small detector to answer that, and the interesting part was not the code — it was watching where it confidently fooled itself.

Here is the whole thing, runnable as-is.

``` python
import re
from collections import Counter

def sentences(text):
    # split on . ! ? followed by whitespace; keeps the delimiter
    return [s.strip() for s in re.split(r'(?<=[.!?])\s+', text) if s.strip()]

def burstiness(text):
    lens = [len(s.split()) for s in sentences(text)]
    if len(lens) < 3:
        return None  # too short to mean anything
    mean = sum(lens) / len(lens)
    var = sum((x - mean) ** 2 for x in lens) / len(lens)
    return var ** 0.5 / mean  # coefficient of variation

def repeated_ngrams(text, n=3, threshold=2):
    words = re.findall(r"[a-z']+", text.lower())
    grams = [" ".join(words[i:i+n]) for i in range(len(words)-n+1)]
    return {g: c for g, c in Counter(grams).items() if c >= threshold}

def score(text):
    b = burstiness(text)
    rep = repeated_ngrams(text)
    # low burstiness + repeated phrases => suspicious
    if b is None:
        return None
    return round((1 - min(b, 1)) * 0.6 + min(len(rep), 10) / 10 * 0.4, 3)

if __name__ == "__main__":
    human = "I shipped the proxy on a Tuesday. Two days later the stream broke. Turned out the client buffered wrong. Fixed it by Friday."
    robot = "The implementation of the proxy was completed successfully. The system was designed to be robust. The architecture ensures reliability. The solution provides scalability. The approach demonstrates effectiveness."
    print("human:", score(human))   # usually higher (varied lengths)
    print("robot:", score(robot))   # usually lower (uniform, repetitive)
```

Run it and you will see the robot sample scores lower almost every time. Good enough to sort a pile of drafts, right? Three things broke that assumption in production.

**Pitfall 1: short texts are noise.** `burstiness` returns `None` under 3 sentences, and even at 3–4 sentences the coefficient of variation swings wildly. A two-sentence human reply can look more "AI" than a ten-sentence marketing page. I had to treat anything under ~40 words as "unknown" instead of scoring it.

**Pitfall 2: genre beats author.** Legal summaries, release notes, and API docs are supposed to be uniform. My detector flagged real human technical writing as AI because the domain is naturally low-burstiness. Burstiness alone is a genre signal, not an authorship signal.

**Pitfall 3: the detector is the easiest thing to game.** Swapping two sentences, splitting one long one into two, or changing "utilize" to "use" drops the repeated-ngram count to zero. That is exactly what a cheap humanizer does — it moves the score without moving the prose.

So detection helped me triage, but it never told me the text was *good*. The part that actually changed how the output reads was a fixed set of rewrite moves: vary sentence openings, cut the second "that", replace noun-stacks with verbs, and break the rhythm every third sentence. We baked those into a tool we use daily — [ShipCopy](https://launchcraft.io/shipcopy) — but the patterns work fine by hand once you have seen them fail a detector a few times.

The takeaway: a 40-line detector is a fine smoke test. It is not a judge of quality, and anyone who promises the text is human is selling the score, not the sentence.
