I built Podium, a tool that compares your reading of a speech with a great delivery of the same text (JFK, Reagan) and tells you exactly where you drift and why: "14.3 syllables/s here vs 7.7 in the reference (+85%)", pinned to a time range.
The first version worked beautifully on my test set. Then I gave it a different speaker, and it flagged 41 "flaws" per minute on a perfectly clean recording.
This post covers why that happened, the one-line fix, and the evaluation set-up that caught it.
Comparing two deliveries frame by frame is hopeless: they never line up. Comparing them word by word is easy, as long as both are force-aligned to the same transcript. I used torchaudio's MMS_FA (a wav2vec2 CTC aligner):
bundle = torchaudio.pipelines.MMS_FA
model, tokenizer, aligner = bundle.get_model(with_star=False), bundle.get_tokenizer(), bundle.get_aligner()
emission, _ = model(wav)
spans = aligner(emission[0], tokenizer(words)) # one span list per word
Now word 17 of your reading is word 17 of JFK's. For every word I compute speaker-normalised features with Praat (via Parselmouth) and librosa:
There's no public dataset that pairs a good delivery with bad deliveries of the same words. So I made one: 207 recordings, starting from three public-domain speeches.
The trick is that I inject the flaws myself, so every label is sample-accurate:
| Flaw | How it's injected | Severity 1 / 2 / 3 |
|---|---|---|
| rushed / dragged | phase-vocoder time-scale of 4–9 words | ×1.3/1.6/2.0 · ×0.8/0.65/0.5 |
| monotone | WORLD vocoder, pitch contour squashed toward its mean | 50/75/95% removed |
| mumbled | gain down + low-pass | −6 dB @ 3 kHz … −16 dB @ 1.1 kHz |
| awkward | room-tone silence inserted mid-phrase | 0.7 / 1.3 / 2.2 s |
| missing breath | a natural squeezed out | 50 / 20 / 0% kept |
| stutter | word onset repeated | 1 / 2 / 3 repeats |
Edits are joined with 5 ms fades that preserve length, so the label times never drift. My first version used 15 ms crossfades, which quietly shifted every later label by 15 ms per edit.
I also had the same texts read by open Piper TTS voices, with flaws injected into those too. That's the "different speaker" test.
And I split it honestly: thresholds are tuned only on JFK, then frozen and tested on Reagan plus an unseen voice.
Version 1 scored each word by how far it departed from the reference, as a robust z-score (median/MAD) after removing your overall offset. On Reagan's own audio with injected flaws it got F1 0.64 with zero false alarms.
On a TTS voice reading the same text:
| F1 | False alarms on clean audio | |
|---|---|---|
| Same speaker | 0.64 | 0 / min |
| Different speaker | 0.03 | 41 / min |
A synthetic voice doesn't breathe where Reagan breathes or lift the words he lifts. Measured against Reagan, everything it does is a deviation. Technically correct, but useless as feedback: nobody wants to hear that every sentence is wrong because they aren't Reagan.
The fix was to separate style from flaws. A flaw is something that's off compared with the reference and off compared with your own delivery. So every word gets two scores:
z_ref = robust_z(participant - reference) # departs from the reference, beyond your overall style
z_self = robust_z(participant_feature) # stands out within your own recording
flaw = np.fmin(z_ref, z_self) # a soft AND
A consistently different voice moves z_ref everywhere but z_self almost nowhere, so it's reported once as style ("33% faster and flatter than the reference") instead of 40 times as flaws.
False alarms on clean different-speaker recordings went from 41/min to 3/min.
Injected stutters kept being reported as "awkward ". The cause: the aligner usually places the word at its last restart, so "w- w- we" becomes a long gap followed by a normal "we".
The fix was one feature: seconds of voiced audio inside the gap. A real is silent; a gap full of voiced sound is a restart.
gap_voiced = np.sum(~np.isnan(f0[prev_end:word_start])) * HOP
Long s now require gap_voiced < 0.15 s, and stutters can trigger on gap_voiced > 0.12 s. Stutter recall went from 0.11 to 0.44 on the held-out set.
What doesn't work yet:
Code, dataset (207 labelled recordings) and the full evaluation are open source:
Read a great speech. Podium shows exactly where your delivery drifts from it, to the tenth of a second and explains each flaw with the numbers behind it.
Built for the Multimodal AI Hackathon 2026, Track C: Contrastive Speech Analytics & Temporal Flaw Grounding.
Built for the Multimodal AI Hackathon 2026 (Track C), with heavy help from Claude Code, an AI coding agent, for implementation and evaluation runs.