# My speech-flaw detector flagged 41 false alarms a minute. One line of math fixed it.

> Source: <https://dev.to/jaypokale/my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed-it-1id1>
> Published: 2026-10-06 20:04:50+00:00

I built **Podium**, a tool that compares your reading of a speech with a great delivery of the *same text* (JFK, Reagan) and tells you exactly where you drift and why: "14.3 syllables/s here vs 7.7 in the reference (+85%)", pinned to a time range.

The first version worked beautifully on my test set. Then I gave it a different speaker, and it flagged **41 "flaws" per minute** on a perfectly clean recording.

This post covers why that happened, the one-line fix, and the evaluation set-up that caught it.

Comparing two deliveries frame by frame is hopeless: they never line up. Comparing them *word by word* is easy, as long as both are force-aligned to the same transcript. I used torchaudio's `MMS_FA` (a wav2vec2 CTC aligner):

```
bundle = torchaudio.pipelines.MMS_FA
model, tokenizer, aligner = bundle.get_model(with_star=False), bundle.get_tokenizer(), bundle.get_aligner()
emission, _ = model(wav)
spans = aligner(emission[0], tokenizer(words))   # one span list per word
```

Now word 17 of your reading is word 17 of JFK's. For every word I compute speaker-normalised features with Praat (via Parselmouth) and librosa:

There's no public dataset that pairs a good delivery with bad deliveries of the same words. So I made one: **207 recordings**, starting from three public-domain speeches.

The trick is that I *inject* the flaws myself, so every label is sample-accurate:

| Flaw | How it's injected | Severity 1 / 2 / 3 | 
|---|---|---|
| rushed / dragged | phase-vocoder time-scale of 4–9 words | ×1.3/1.6/2.0 · ×0.8/0.65/0.5 | 
| monotone | WORLD vocoder, pitch contour squashed toward its mean | 50/75/95% removed | 
| mumbled | gain down + low-pass | −6 dB @ 3 kHz … −16 dB @ 1.1 kHz | 
| awkward pause | room-tone silence inserted mid-phrase | 0.7 / 1.3 / 2.2 s | 
| missing breath | a natural pause squeezed out | 50 / 20 / 0% kept | 
| stutter | word onset repeated | 1 / 2 / 3 repeats | 

Edits are joined with 5 ms fades that **preserve length**, so the label times never drift. My first version used 15 ms crossfades, which quietly shifted every later label by 15 ms per edit.

I also had the same texts read by open **Piper** TTS voices, with flaws injected into those too. That's the "different speaker" test.

And I split it honestly: thresholds are tuned **only on JFK**, then frozen and tested on **Reagan** plus an unseen voice.

Version 1 scored each word by how far it departed from the reference, as a robust z-score (median/MAD) after removing your overall offset. On Reagan's own audio with injected flaws it got F1 0.64 with **zero** false alarms.

On a TTS voice reading the same text:

|  | F1 | False alarms on clean audio | 
|---|---|---|
| Same speaker | 0.64 | 0 / min | 
| **Different speaker** | **0.03** | **41 / min** | 

A synthetic voice doesn't breathe where Reagan breathes or lift the words he lifts. Measured against Reagan, *everything* it does is a deviation. Technically correct, but useless as feedback: nobody wants to hear that every sentence is wrong because they aren't Reagan.

The fix was to separate **style** from **flaws**. A flaw is something that's off compared with the reference *and* off compared with **your own** delivery. So every word gets two scores:

```
z_ref  = robust_z(participant - reference)   # departs from the reference, beyond your overall style
z_self = robust_z(participant_feature)       # stands out within your own recording
flaw   = np.fmin(z_ref, z_self)              # a soft AND
```

A consistently different voice moves `z_ref` everywhere but `z_self` almost nowhere, so it's reported once as *style* ("33% faster and flatter than the reference") instead of 40 times as flaws.

**False alarms on clean different-speaker recordings went from 41/min to 3/min.**

Injected stutters kept being reported as "awkward pause". The cause: the aligner usually places the word at its *last* restart, so "w- w- we" becomes a long gap followed by a normal "we".

The fix was one feature: **seconds of voiced audio inside the gap**. A real pause is silent; a gap full of voiced sound is a restart.

```
gap_voiced = np.sum(~np.isnan(f0[prev_end:word_start])) * HOP
```

Long pauses now require `gap_voiced < 0.15 s`, and stutters can trigger on `gap_voiced > 0.12 s`. Stutter recall went from 0.11 to 0.44 on the held-out set.

What doesn't work yet:

Code, dataset (207 labelled recordings) and the full evaluation are open source:

**Read a great speech. Podium shows exactly where your delivery drifts from it, to the tenth of a second
and explains each flaw with the numbers behind it.**

Built for the Multimodal AI Hackathon 2026, **Track C: Contrastive Speech Analytics & Temporal Flaw Grounding**.

*Built for the Multimodal AI Hackathon 2026 (Track C), with heavy help from Claude Code, an AI coding agent, for implementation and evaluation runs.*
