STT model recommendations Hugging Face forum users recommended WhisperX, which combines Whisper large-v3 or large-v3-turbo transcription with Pyannote 3.1 diarization, for transcribing a one-hour two-person English podcast with casual overlapping speech. WhisperX performs forced alignment after transcription rather than in strict real-time order, so interjections like "Mmm, right, that makes sen-" are attributed correctly; large-v3-turbo is 4x faster than large-v3 with barely worse accuracy. A commenter identifying as the developer of skeletonfingers.com noted its fast option labels speakers and is free for files up to 30 minutes (3 per day), with a 1-hour episode requiring a $5 pack or splitting into two files, while its unlimited in-browser mode runs Whisper locally without speaker labels. Jtxhob https://discuss.huggingface.co/u/Jtxhob 1 Hi all, I am looking for a decent STT model that can transcribe audio from a podcast. I’ve listed some details about the content of the audio file, in case that helps: - Straightforward, English-speaking, two-person-conversation audio that is approximately 1 hour in length. - The conversation between the podcast speakers is casual, so it does contain typical filled pauses, discourse markers, disfluencies, etc. - Timestamps aren’t necessary. - Identifying the active speaker and what they have generally said is necessary, but the precision with which the transcript reports this is not. - For example, suppose Jack and Jill are the speakers. Jill is explaining something, and while she is speaking, Jack interjects with a filler line like “Mmm, right, that makes sen-”, only for Jack to pause so that Jill’s initial speaking line can continue uninterrupted. If the STT were to transcribe this precisely, it would need to very suddenly realise that Jill is no longer speaking, create a new speaker line for Jack, only to return back to Jill again. I can see lots of issues with that in a more casual and non-turn-based conversation, such as inaccuracies of who said what exactly, and in distinguishing a speaker’s intent with what is literally said. WhisperX is the fit for what you’re describing, it bolts word-level alignment plus diarization via Pyannote onto Whisper, and critically it does forced-alignment after transcription rather than transcribing interruptions in strict real-time order, so those messy overlapping-speaker moments “Mmm, right, that makes sen-” get attributed correctly without needing frame-perfect turn-taking detection. Since you don’t need precise timestamps, you can skip the alignment step and just use WhisperX’s diarization pipeline on top of Whisper large-v3 or large-v3-turbo turbo is 4x faster with barely worse accuracy, worth it for a 1hr file . Plain Whisper alone no WhisperX has no built-in speaker separation, so you’d get a wall of text with no idea who said what, that’s the actual gap you’d hit for a two-person podcast without adding diarization on top. Pyannote 3.1 is the diarization backbone WhisperX uses and is currently the best-supported open option for exactly this “good enough, not perfect” speaker separation on casual conversational audio. Agree WhisperX makes sense for this. If you have a gpu you can run large-v3-turbo. If you don’t want to set it up yourself you can use https://skeletonfingers.com https://skeletonfingers.com disclaimer: I’m the dev behind it . The fast option labels speakers. It’s free for files up to 30min 3 a day , so a 1-hour episode needs a $5 pack or splitting into 2 files. There’s also an unlimited free in-browser mode that runs Whisper on your machine, but it doesn’t label speakers.