Building a Private, Offline Video Transcription Pipeline with Whisper AI A developer built Video Transcriber Pro, a desktop tool that transcribes video and audio locally using OpenAI's Whisper model, ensuring sensitive material never leaves the machine. The pipeline integrates yt-dlp, ffmpeg, pyannote.audio for speaker diarization, and optional cloud services like Deepgram and Claude, offering a private alternative to cloud-based transcription. The Problem with Cloud Transcription Most transcription tools today route your audio through someone else's server. That means: For journalists, researchers, and content creators working with sensitive material, that's a non-starter. A confidential interview shouldn't pass through a third-party API before it reaches your transcript. The alternative is to run speech-to-text locally. With OpenAI's Whisper model open-sourced under MIT, that's now practical on a mid-range laptop. What We're Building A Windows desktop workflow that: This is the architecture behind Video Transcriber Pro https://store.rubibot.org/l/videotranscriberpro , a desktop tool I built to scratch my own itch — I needed to transcribe long research interviews without uploading them anywhere. The Stack | Component | Choice | Why | |---|---|---| | Speech-to-text local | Whisper openai-whisper | MIT-licensed, runs on CPU or GPU | | Speech-to-text cloud | Deepgram Nova-2 | Fast, accurate, pay-per-use | | Audio source | yt-dlp + ffmpeg | Pulls audio from YouTube, decodes files | | Speaker diarization | pyannote.audio | Open-source, speaker segmentation | | Text polish | Claude optional | Grammar, fluency, translation | | Export | python-docx, srt | DOCX + SRT + plain text | Step 1: Pull Audio from a YouTube URL The first hurdle is getting clean audio out of a YouTube video. yt-dlp handles this in a few lines: python import subprocess def download youtube audio url: str, output path: str = "audio.mp3" - str: subprocess.run "yt-dlp", "-x", extract audio only "--audio-format", "mp3", "--audio-quality", "0", best "-o", output path, url , check=True return output path For local files, ffmpeg does the conversion to a Whisper-friendly format 16 kHz mono WAV : php def to wav input path: str, output path: str = "audio.wav" - str: subprocess.run "ffmpeg", "-y", "-i", input path, "-ar", "16000", 16 kHz "-ac", "1", mono "-c:a", "pcm s16le", output path , check=True return output path The 16 kHz mono downmix matters. Whisper was trained on this format, and feeding it 48 kHz stereo audio noticeably degrades accuracy. Step 2: Run Whisper Locally The simplest path is the whisper Python package: python import whisper model = whisper.load model "medium" base, small, medium, large result = model.transcribe "audio.wav", language="en" print result "text" For long files, you'll want to stream segments instead of holding the whole transcript in memory: for segment in model.transcribe "audio.wav" : print f" {segment 'start' :.1f}s {segment 'text' }" Model size tradeoff: | Model | Size | Speed CPU | Accuracy | |---|---|---|---| | tiny | 39M | Very fast | Low | | base | 74M | Fast | Medium | | small | 244M | Medium | Good | | medium | 769M | Slow | High | | large | 1550M | Very slow | Best | On a laptop with an RTX 4050, medium runs roughly 4x real-time on GPU. On CPU, expect closer to 0.3x real-time — fine for short clips, painful for 2-hour interviews. Step 3: Speaker Diarization Whisper gives you the words, but not who said them. pyannote.audio fills that gap: python from pyannote.audio import Pipeline pipeline = Pipeline.from pretrained "pyannote/speaker-diarization-3.1", use auth token="YOUR HF TOKEN" diarization = pipeline "audio.wav", min speakers=2, max speakers=5 for turn, , speaker in diarization.itertracks yield label=True : print f"{turn.start:.1f}-{turn.end:.1f}s: {speaker}" You then merge Whisper's segments with pyannote's speaker turns by timestamp overlap. The result is a transcript where each line is attributed: 00:12.3 - 00:14.1 SPEAKER 00: So what made you start this project? 00:14.5 - 00:19.8 SPEAKER 01: Mostly frustration with existing tools. Step 4: Export to SRT Subtitles SRT is just a text format with timestamps. Generating it from Whisper segments is straightforward: python def to srt segments, path: str : def fmt t : h = int t // 3600 m = int t % 3600 // 60 s = int t % 60 ms = int t - int t 1000 return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}" with open path, "w", encoding="utf-8" as f: for i, seg in enumerate segments, 1 : f.write f"{i}\n" f.write f"{fmt seg 'start' } -- {fmt seg 'end' }\n" f.write f"{seg 'text' .strip }\n\n" For DOCX, python-docx writes the same content into a Word document with headings per speaker. Step 5: Optional Cloud Mode Deepgram Nova-2 When privacy isn't a concern and you need speed, Deepgram's Nova-2 is hard to beat. The API is a single WebSocket call: python from deepgram import DeepgramClient dg = DeepgramClient "YOUR DEEPGRAM KEY" options = { "model": "nova-2", "smart format": True, "diarize": True, "utterances": True, } response = dg.listen.prerecorded.v "1" .transcribe file {"buffer": audio bytes}, options Nova-2 typically returns a 30-minute file in under 30 seconds. Whisper medium on GPU takes ~8 minutes for the same file. The tradeoff is clear: cloud for speed, local for privacy. The Privacy Argument This is the part that matters most for a lot of users: For a journalist working with a confidential source, or a researcher handling IRB-protected interviews, only local mode is acceptable. That's why the tool defaults to Whisper and treats the cloud path as opt-in. Where This Lives I packaged this pipeline into a Windows desktop app so non-technical users can run it without touching a terminal: It's called Video Transcriber Pro https://store.rubibot.org/l/videotranscriberpro and it's a one-time purchase — no subscription, no per-minute billing. The local mode requires no API key at all. If you just want the code, every piece above is open-source and composable. If you want the polished desktop version, that's the product. FAQ Does local mode really need no internet? After the first run which downloads the Whisper model weights, ~1.5 GB for medium , yes — fully offline. How accurate is Whisper medium ? On clean English audio, word error rate is typically 5–8%. Noisy recordings and heavy accents degrade it, but that's true of every ASR system. Can I run this on a Mac or Linux? The pipeline is pure Python and works cross-platform. The desktop app I linked is Windows-only because that's where most of my users are, but the underlying code runs anywhere ffmpeg and whisper do. What about real-time transcription? Whisper supports streaming via whisper-streaming or the faster-whisper backend. The desktop app currently does batch transcription; real-time is on the roadmap. Takeaway Speech-to-text is a solved problem technically. The open question is whether you're willing to upload your audio to solve it. With Whisper running locally, you don't have to — and you still get accuracy good enough for subtitles, meeting notes, and interview transcripts. The code is open. The packaged app is product here https://store.rubibot.org/l/videotranscriberpro if you want the desktop version. If you found this useful, I write about AI tooling and agent architectures at dev.to/devhunterai https://dev.to/devhunterai .