cd /news/machine-learning/my-speech-flaw-detector-flagged-41-f… · home › topics › machine-learning › article
[ARTICLE · art-146335] src=dev.to ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

My speech-flaw detector flagged 41 false alarms a minute. One line of math fixed it.

A developer built Podium, a speech-delivery analysis tool that force-aligns a speaker's reading of a public-domain speech (JFK, Reagan) word-by-word against a reference delivery using torchaudio's MMS_FA wav2vec2 CTC aligner, then scores each word with Praat and librosa features. Version 1 flagged 41 false alarms per minute on a clean recording from a different speaker (F1 0.03), which the developer fixed by combining a reference-relative z-score with a within-speaker z-score via np.fmin, cutting false alarms to 3 per minute and separating persistent voice style from genuine flaws.

by read4 min views1 publishedOct 6, 2026

I built Podium, a tool that compares your reading of a speech with a great delivery of the same text (JFK, Reagan) and tells you exactly where you drift and why: "14.3 syllables/s here vs 7.7 in the reference (+85%)", pinned to a time range.

The first version worked beautifully on my test set. Then I gave it a different speaker, and it flagged 41 "flaws" per minute on a perfectly clean recording.

This post covers why that happened, the one-line fix, and the evaluation set-up that caught it.

Comparing two deliveries frame by frame is hopeless: they never line up. Comparing them word by word is easy, as long as both are force-aligned to the same transcript. I used torchaudio's MMS_FA (a wav2vec2 CTC aligner):

bundle = torchaudio.pipelines.MMS_FA
model, tokenizer, aligner = bundle.get_model(with_star=False), bundle.get_tokenizer(), bundle.get_aligner()
emission, _ = model(wav)
spans = aligner(emission[0], tokenizer(words))   # one span list per word

Now word 17 of your reading is word 17 of JFK's. For every word I compute speaker-normalised features with Praat (via Parselmouth) and librosa:

There's no public dataset that pairs a good delivery with bad deliveries of the same words. So I made one: 207 recordings, starting from three public-domain speeches.

The trick is that I inject the flaws myself, so every label is sample-accurate:

Flaw How it's injected Severity 1 / 2 / 3
rushed / dragged phase-vocoder time-scale of 4–9 words ×1.3/1.6/2.0 · ×0.8/0.65/0.5
monotone WORLD vocoder, pitch contour squashed toward its mean 50/75/95% removed
mumbled gain down + low-pass −6 dB @ 3 kHz … −16 dB @ 1.1 kHz
awkward room-tone silence inserted mid-phrase 0.7 / 1.3 / 2.2 s
missing breath a natural squeezed out 50 / 20 / 0% kept
stutter word onset repeated 1 / 2 / 3 repeats

Edits are joined with 5 ms fades that preserve length, so the label times never drift. My first version used 15 ms crossfades, which quietly shifted every later label by 15 ms per edit.

I also had the same texts read by open Piper TTS voices, with flaws injected into those too. That's the "different speaker" test.

And I split it honestly: thresholds are tuned only on JFK, then frozen and tested on Reagan plus an unseen voice.

Version 1 scored each word by how far it departed from the reference, as a robust z-score (median/MAD) after removing your overall offset. On Reagan's own audio with injected flaws it got F1 0.64 with zero false alarms.

On a TTS voice reading the same text:

F1 False alarms on clean audio
Same speaker 0.64 0 / min
Different speaker 0.03 41 / min

A synthetic voice doesn't breathe where Reagan breathes or lift the words he lifts. Measured against Reagan, everything it does is a deviation. Technically correct, but useless as feedback: nobody wants to hear that every sentence is wrong because they aren't Reagan.

The fix was to separate style from flaws. A flaw is something that's off compared with the reference and off compared with your own delivery. So every word gets two scores:

z_ref  = robust_z(participant - reference)   # departs from the reference, beyond your overall style
z_self = robust_z(participant_feature)       # stands out within your own recording
flaw   = np.fmin(z_ref, z_self)              # a soft AND

A consistently different voice moves z_ref everywhere but z_self almost nowhere, so it's reported once as style ("33% faster and flatter than the reference") instead of 40 times as flaws.

False alarms on clean different-speaker recordings went from 41/min to 3/min.

Injected stutters kept being reported as "awkward ". The cause: the aligner usually places the word at its last restart, so "w- w- we" becomes a long gap followed by a normal "we".

The fix was one feature: seconds of voiced audio inside the gap. A real is silent; a gap full of voiced sound is a restart.

gap_voiced = np.sum(~np.isnan(f0[prev_end:word_start])) * HOP

Long s now require gap_voiced < 0.15 s, and stutters can trigger on gap_voiced > 0.12 s. Stutter recall went from 0.11 to 0.44 on the held-out set.

What doesn't work yet:

Code, dataset (207 labelled recordings) and the full evaluation are open source:

Read a great speech. Podium shows exactly where your delivery drifts from it, to the tenth of a second and explains each flaw with the numbers behind it.

Built for the Multimodal AI Hackathon 2026, Track C: Contrastive Speech Analytics & Temporal Flaw Grounding.

Built for the Multimodal AI Hackathon 2026 (Track C), with heavy help from Claude Code, an AI coding agent, for implementation and evaluation runs.

── more in #machine-learning 4 stories · sorted by recency
── more on @podium 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/my-speech-flaw-detec…] indexed:0 read:4min 2026-10-06 · —