My speech-flaw detector flagged 41 false alarms a minute. One line of math fixed it. A developer built Podium, a speech-delivery analysis tool that force-aligns a speaker's reading of a public-domain speech (JFK, Reagan) word-by-word against a reference delivery using torchaudio's MMS_FA wav2vec2 CTC aligner, then scores each word with Praat and librosa features. Version 1 flagged 41 false alarms per minute on a clean recording from a different speaker (F1 0.03), which the developer fixed by combining a reference-relative z-score with a within-speaker z-score via np.fmin, cutting false alarms to 3 per minute and separating persistent voice style from genuine flaws. I built Podium , a tool that compares your reading of a speech with a great delivery of the same text JFK, Reagan and tells you exactly where you drift and why: "14.3 syllables/s here vs 7.7 in the reference +85% ", pinned to a time range. The first version worked beautifully on my test set. Then I gave it a different speaker, and it flagged 41 "flaws" per minute on a perfectly clean recording. This post covers why that happened, the one-line fix, and the evaluation set-up that caught it. Comparing two deliveries frame by frame is hopeless: they never line up. Comparing them word by word is easy, as long as both are force-aligned to the same transcript. I used torchaudio's MMS FA a wav2vec2 CTC aligner : bundle = torchaudio.pipelines.MMS FA model, tokenizer, aligner = bundle.get model with star=False , bundle.get tokenizer , bundle.get aligner emission, = model wav spans = aligner emission 0 , tokenizer words one span list per word Now word 17 of your reading is word 17 of JFK's. For every word I compute speaker-normalised features with Praat via Parselmouth and librosa: There's no public dataset that pairs a good delivery with bad deliveries of the same words. So I made one: 207 recordings , starting from three public-domain speeches. The trick is that I inject the flaws myself, so every label is sample-accurate: | Flaw | How it's injected | Severity 1 / 2 / 3 | |---|---|---| | rushed / dragged | phase-vocoder time-scale of 4–9 words | ×1.3/1.6/2.0 · ×0.8/0.65/0.5 | | monotone | WORLD vocoder, pitch contour squashed toward its mean | 50/75/95% removed | | mumbled | gain down + low-pass | −6 dB @ 3 kHz … −16 dB @ 1.1 kHz | | awkward pause | room-tone silence inserted mid-phrase | 0.7 / 1.3 / 2.2 s | | missing breath | a natural pause squeezed out | 50 / 20 / 0% kept | | stutter | word onset repeated | 1 / 2 / 3 repeats | Edits are joined with 5 ms fades that preserve length , so the label times never drift. My first version used 15 ms crossfades, which quietly shifted every later label by 15 ms per edit. I also had the same texts read by open Piper TTS voices, with flaws injected into those too. That's the "different speaker" test. And I split it honestly: thresholds are tuned only on JFK , then frozen and tested on Reagan plus an unseen voice. Version 1 scored each word by how far it departed from the reference, as a robust z-score median/MAD after removing your overall offset. On Reagan's own audio with injected flaws it got F1 0.64 with zero false alarms. On a TTS voice reading the same text: | | F1 | False alarms on clean audio | |---|---|---| | Same speaker | 0.64 | 0 / min | | Different speaker | 0.03 | 41 / min | A synthetic voice doesn't breathe where Reagan breathes or lift the words he lifts. Measured against Reagan, everything it does is a deviation. Technically correct, but useless as feedback: nobody wants to hear that every sentence is wrong because they aren't Reagan. The fix was to separate style from flaws . A flaw is something that's off compared with the reference and off compared with your own delivery. So every word gets two scores: z ref = robust z participant - reference departs from the reference, beyond your overall style z self = robust z participant feature stands out within your own recording flaw = np.fmin z ref, z self a soft AND A consistently different voice moves z ref everywhere but z self almost nowhere, so it's reported once as style "33% faster and flatter than the reference" instead of 40 times as flaws. False alarms on clean different-speaker recordings went from 41/min to 3/min. Injected stutters kept being reported as "awkward pause". The cause: the aligner usually places the word at its last restart, so "w- w- we" becomes a long gap followed by a normal "we". The fix was one feature: seconds of voiced audio inside the gap . A real pause is silent; a gap full of voiced sound is a restart. gap voiced = np.sum ~np.isnan f0 prev end:word start HOP Long pauses now require gap voiced < 0.15 s , and stutters can trigger on gap voiced 0.12 s . Stutter recall went from 0.11 to 0.44 on the held-out set. What doesn't work yet: Code, dataset 207 labelled recordings and the full evaluation are open source: Read a great speech. Podium shows exactly where your delivery drifts from it, to the tenth of a second and explains each flaw with the numbers behind it. Built for the Multimodal AI Hackathon 2026, Track C: Contrastive Speech Analytics & Temporal Flaw Grounding . Built for the Multimodal AI Hackathon 2026 Track C , with heavy help from Claude Code, an AI coding agent, for implementation and evaluation runs.