{"slug": "my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed", "title": "My speech-flaw detector flagged 41 false alarms a minute. One line of math fixed it.", "summary": "A developer built Podium, a speech-delivery analysis tool that force-aligns a speaker's reading of a public-domain speech (JFK, Reagan) word-by-word against a reference delivery using torchaudio's MMS_FA wav2vec2 CTC aligner, then scores each word with Praat and librosa features. Version 1 flagged 41 false alarms per minute on a clean recording from a different speaker (F1 0.03), which the developer fixed by combining a reference-relative z-score with a within-speaker z-score via np.fmin, cutting false alarms to 3 per minute and separating persistent voice style from genuine flaws.", "body_md": "I built **Podium**, a tool that compares your reading of a speech with a great delivery of the *same text* (JFK, Reagan) and tells you exactly where you drift and why: \"14.3 syllables/s here vs 7.7 in the reference (+85%)\", pinned to a time range.\n\nThe first version worked beautifully on my test set. Then I gave it a different speaker, and it flagged **41 \"flaws\" per minute** on a perfectly clean recording.\n\nThis post covers why that happened, the one-line fix, and the evaluation set-up that caught it.\n\nComparing two deliveries frame by frame is hopeless: they never line up. Comparing them *word by word* is easy, as long as both are force-aligned to the same transcript. I used torchaudio's `MMS_FA` (a wav2vec2 CTC aligner):\n\n```\nbundle = torchaudio.pipelines.MMS_FA\nmodel, tokenizer, aligner = bundle.get_model(with_star=False), bundle.get_tokenizer(), bundle.get_aligner()\nemission, _ = model(wav)\nspans = aligner(emission[0], tokenizer(words))   # one span list per word\n```\n\nNow word 17 of your reading is word 17 of JFK's. For every word I compute speaker-normalised features with Praat (via Parselmouth) and librosa:\n\nThere's no public dataset that pairs a good delivery with bad deliveries of the same words. So I made one: **207 recordings**, starting from three public-domain speeches.\n\nThe trick is that I *inject* the flaws myself, so every label is sample-accurate:\n\n| Flaw | How it's injected | Severity 1 / 2 / 3 | \n|---|---|---|\n| rushed / dragged | phase-vocoder time-scale of 4–9 words | ×1.3/1.6/2.0 · ×0.8/0.65/0.5 | \n| monotone | WORLD vocoder, pitch contour squashed toward its mean | 50/75/95% removed | \n| mumbled | gain down + low-pass | −6 dB @ 3 kHz … −16 dB @ 1.1 kHz | \n| awkward pause | room-tone silence inserted mid-phrase | 0.7 / 1.3 / 2.2 s | \n| missing breath | a natural pause squeezed out | 50 / 20 / 0% kept | \n| stutter | word onset repeated | 1 / 2 / 3 repeats | \n\nEdits are joined with 5 ms fades that **preserve length**, so the label times never drift. My first version used 15 ms crossfades, which quietly shifted every later label by 15 ms per edit.\n\nI also had the same texts read by open **Piper** TTS voices, with flaws injected into those too. That's the \"different speaker\" test.\n\nAnd I split it honestly: thresholds are tuned **only on JFK**, then frozen and tested on **Reagan** plus an unseen voice.\n\nVersion 1 scored each word by how far it departed from the reference, as a robust z-score (median/MAD) after removing your overall offset. On Reagan's own audio with injected flaws it got F1 0.64 with **zero** false alarms.\n\nOn a TTS voice reading the same text:\n\n|  | F1 | False alarms on clean audio | \n|---|---|---|\n| Same speaker | 0.64 | 0 / min | \n| **Different speaker** | **0.03** | **41 / min** | \n\nA synthetic voice doesn't breathe where Reagan breathes or lift the words he lifts. Measured against Reagan, *everything* it does is a deviation. Technically correct, but useless as feedback: nobody wants to hear that every sentence is wrong because they aren't Reagan.\n\nThe fix was to separate **style** from **flaws**. A flaw is something that's off compared with the reference *and* off compared with **your own** delivery. So every word gets two scores:\n\n```\nz_ref  = robust_z(participant - reference)   # departs from the reference, beyond your overall style\nz_self = robust_z(participant_feature)       # stands out within your own recording\nflaw   = np.fmin(z_ref, z_self)              # a soft AND\n```\n\nA consistently different voice moves `z_ref` everywhere but `z_self` almost nowhere, so it's reported once as *style* (\"33% faster and flatter than the reference\") instead of 40 times as flaws.\n\n**False alarms on clean different-speaker recordings went from 41/min to 3/min.**\n\nInjected stutters kept being reported as \"awkward pause\". The cause: the aligner usually places the word at its *last* restart, so \"w- w- we\" becomes a long gap followed by a normal \"we\".\n\nThe fix was one feature: **seconds of voiced audio inside the gap**. A real pause is silent; a gap full of voiced sound is a restart.\n\n```\ngap_voiced = np.sum(~np.isnan(f0[prev_end:word_start])) * HOP\n```\n\nLong pauses now require `gap_voiced < 0.15 s`, and stutters can trigger on `gap_voiced > 0.12 s`. Stutter recall went from 0.11 to 0.44 on the held-out set.\n\nWhat doesn't work yet:\n\nCode, dataset (207 labelled recordings) and the full evaluation are open source:\n\n**Read a great speech. Podium shows exactly where your delivery drifts from it, to the tenth of a second\nand explains each flaw with the numbers behind it.**\n\nBuilt for the Multimodal AI Hackathon 2026, **Track C: Contrastive Speech Analytics & Temporal Flaw Grounding**.\n\n*Built for the Multimodal AI Hackathon 2026 (Track C), with heavy help from Claude Code, an AI coding agent, for implementation and evaluation runs.*", "url": "https://wpnews.pro/news/my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed", "canonical_source": "https://dev.to/jaypokale/my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed-it-1id1", "published_at": "2026-10-06 20:04:50+00:00", "updated_at": "2026-10-06 20:18:33.651506+00:00", "lang": "en", "topics": ["machine-learning", "natural-language-processing", "ai-tools"], "entities": ["Podium", "torchaudio", "MMS_FA", "Praat", "Parselmouth", "librosa", "Piper", "WORLD vocoder"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed", "markdown": "https://wpnews.pro/news/my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed.md", "text": "https://wpnews.pro/news/my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed.txt", "jsonld": "https://wpnews.pro/news/my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed.jsonld"}}