cd /news/artificial-intelligence/stt-model-recommendations · home › topics › artificial-intelligence › article
[ARTICLE · art-146706] src=discuss.huggingface.co ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

STT model recommendations

Hugging Face forum users recommended WhisperX, which combines Whisper large-v3 or large-v3-turbo transcription with Pyannote 3.1 diarization, for transcribing a one-hour two-person English podcast with casual overlapping speech. WhisperX performs forced alignment after transcription rather than in strict real-time order, so interjections like "Mmm, right, that makes sen-" are attributed correctly; large-v3-turbo is 4x faster than large-v3 with barely worse accuracy. A commenter identifying as the developer of skeletonfingers.com noted its fast option labels speakers and is free for files up to 30 minutes (3 per day), with a 1-hour episode requiring a $5 pack or splitting into two files, while its unlimited in-browser mode runs Whisper locally without speaker labels.

read2 min views1 publishedOct 6, 2026
STT model recommendations
Image: Discuss (auto-discovered)

Jtxhob 1

Hi all, I am looking for a decent STT model that can transcribe audio from a podcast. I’ve listed some details about the content of the audio file, in case that helps:

  • Straightforward, English-speaking, two-person-conversation audio that is approximately 1 hour in length.
  • The conversation between the podcast speakers is casual, so it does contain typical filled s, discourse markers, disfluencies, etc.
  • Timestamps aren’t necessary.
  • Identifying the active speaker and what they have (generally) said is necessary, but the precision with which the transcript reports this is not.
    • For example, suppose Jack and Jill are the speakers. Jill is explaining something, and while she is speaking, Jack interjects with a filler line like “Mmm, right, that makes sen-”, only for Jack to so that Jill’s initial speaking line can continue uninterrupted. If the STT were to transcribe this precisely, it would need to very suddenly realise that Jill is no longer speaking, create a new speaker line for Jack, only to return back to Jill again. I can see lots of issues with that in a more casual and non-turn-based conversation, such as inaccuracies of who said what exactly, and in distinguishing a speaker’s intent with what is literally said.

WhisperX is the fit for what you’re describing, it bolts word-level alignment plus diarization (via Pyannote) onto Whisper, and critically it does forced-alignment after transcription rather than transcribing interruptions in strict real-time order, so those messy overlapping-speaker moments (“Mmm, right, that makes sen-”) get attributed correctly without needing frame-perfect turn-taking detection. Since you don’t need precise timestamps, you can skip the alignment step and just use WhisperX’s diarization pipeline on top of Whisper large-v3 or large-v3-turbo (turbo is 4x faster with barely worse accuracy, worth it for a 1hr file).

Plain Whisper alone (no WhisperX) has no built-in speaker separation, so you’d get a wall of text with no idea who said what, that’s the actual gap you’d hit for a two-person podcast without adding diarization on top. Pyannote 3.1 is the diarization backbone WhisperX uses and is currently the best-supported open option for exactly this “good enough, not perfect” speaker separation on casual conversational audio.

Agree WhisperX makes sense for this. If you have a gpu you can run large-v3-turbo.

If you don’t want to set it up yourself you can use https://skeletonfingers.com (disclaimer: I’m the dev behind it). The fast option labels speakers. It’s free for files up to 30min (3 a day), so a 1-hour episode needs a $5 pack or splitting into 2 files. There’s also an unlimited free in-browser mode that runs Whisper on your machine, but it doesn’t label speakers.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @hugging face 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stt-model-recommenda…] indexed:0 read:2min 2026-10-06 · —