# Building a Private, Offline Video Transcription Pipeline with Whisper AI

> Source: <https://dev.to/devhunterai/building-a-private-offline-video-transcription-pipeline-with-whisper-ai-3l85>
> Published: 2026-09-09 05:40:44+00:00

The Problem with Cloud Transcription

Most transcription tools today route your audio through someone else's server. That means:

For journalists, researchers, and content creators working with sensitive material, that's a non-starter. A confidential interview shouldn't pass through a third-party API before it reaches your transcript.

The alternative is to run speech-to-text locally. With OpenAI's Whisper model open-sourced under MIT, that's now practical on a mid-range laptop.

What We're Building

A Windows desktop workflow that:

This is the architecture behind [Video Transcriber Pro](https://store.rubibot.org/l/videotranscriberpro), a desktop tool I built to scratch my own itch — I needed to transcribe long research interviews without uploading them anywhere.

The Stack

| Component | Choice | Why | 
|---|---|---|
| Speech-to-text (local) | Whisper (openai-whisper) | MIT-licensed, runs on CPU or GPU | 
| Speech-to-text (cloud) | Deepgram Nova-2 | Fast, accurate, pay-per-use | 
| Audio source | yt-dlp + ffmpeg | Pulls audio from YouTube, decodes files | 
| Speaker diarization | pyannote.audio | Open-source, speaker segmentation | 
| Text polish | Claude (optional) | Grammar, fluency, translation | 
| Export | python-docx, srt | DOCX + SRT + plain text | 

Step 1: Pull Audio from a YouTube URL

The first hurdle is getting clean audio out of a YouTube video. `yt-dlp` handles this in a few lines:

``` python
import subprocess

def download_youtube_audio(url: str, output_path: str = "audio.mp3") -> str:
    subprocess.run([
        "yt-dlp",
        "-x",                       # extract audio only
        "--audio-format", "mp3",
        "--audio-quality", "0",     # best
        "-o", output_path,
        url
    ], check=True)
    return output_path
```

For local files, `ffmpeg` does the conversion to a Whisper-friendly format (16 kHz mono WAV):

``` php
def to_wav(input_path: str, output_path: str = "audio.wav") -> str:
    subprocess.run([
        "ffmpeg", "-y",
        "-i", input_path,
        "-ar", "16000",     # 16 kHz
        "-ac", "1",         # mono
        "-c:a", "pcm_s16le",
        output_path
    ], check=True)
    return output_path
```

The 16 kHz mono downmix matters. Whisper was trained on this format, and feeding it 48 kHz stereo audio noticeably degrades accuracy.

Step 2: Run Whisper Locally

The simplest path is the `whisper` Python package:

``` python
import whisper

model = whisper.load_model("medium")  # base, small, medium, large
result = model.transcribe("audio.wav", language="en")

print(result["text"])
```

For long files, you'll want to stream segments instead of holding the whole transcript in memory:

```
for segment in model.transcribe("audio.wav"):
    print(f"[{segment['start']:.1f}s] {segment['text']}")
```

**Model size tradeoff:**

| Model | Size | Speed (CPU) | Accuracy | 
|---|---|---|---|
| tiny | 39M | Very fast | Low | 
| base | 74M | Fast | Medium | 
| small | 244M | Medium | Good | 
| medium | 769M | Slow | High | 
| large | 1550M | Very slow | Best | 

On a laptop with an RTX 4050, `medium` runs roughly 4x real-time on GPU. On CPU, expect closer to 0.3x real-time — fine for short clips, painful for 2-hour interviews.

Step 3: Speaker Diarization

Whisper gives you the words, but not who said them. `pyannote.audio` fills that gap:

``` python
from pyannote.audio import Pipeline

pipeline = Pipeline.from_pretrained(
    "pyannote/speaker-diarization-3.1",
    use_auth_token="YOUR_HF_TOKEN"
)
diarization = pipeline("audio.wav", min_speakers=2, max_speakers=5)

for turn, _, speaker in diarization.itertracks(yield_label=True):
    print(f"{turn.start:.1f}-{turn.end:.1f}s: {speaker}")
```

You then merge Whisper's segments with pyannote's speaker turns by timestamp overlap. The result is a transcript where each line is attributed:

```
[00:12.3 - 00:14.1] SPEAKER_00: So what made you start this project?
[00:14.5 - 00:19.8] SPEAKER_01: Mostly frustration with existing tools.
```

Step 4: Export to SRT (Subtitles)

SRT is just a text format with timestamps. Generating it from Whisper segments is straightforward:

``` python
def to_srt(segments, path: str):
    def fmt(t):
        h = int(t // 3600)
        m = int((t % 3600) // 60)
        s = int(t % 60)
        ms = int((t - int(t)) * 1000)
        return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}"

    with open(path, "w", encoding="utf-8") as f:
        for i, seg in enumerate(segments, 1):
            f.write(f"{i}\n")
            f.write(f"{fmt(seg['start'])} --> {fmt(seg['end'])}\n")
            f.write(f"{seg['text'].strip()}\n\n")
```

For DOCX, `python-docx` writes the same content into a Word document with headings per speaker.

Step 5: Optional Cloud Mode (Deepgram Nova-2)

When privacy isn't a concern and you need speed, Deepgram's Nova-2 is hard to beat. The API is a single WebSocket call:

``` python
from deepgram import DeepgramClient

dg = DeepgramClient("YOUR_DEEPGRAM_KEY")
options = {
    "model": "nova-2",
    "smart_format": True,
    "diarize": True,
    "utterances": True,
}
response = dg.listen.prerecorded.v("1").transcribe_file({"buffer": audio_bytes}, options)
```

Nova-2 typically returns a 30-minute file in under 30 seconds. Whisper `medium` on GPU takes ~8 minutes for the same file. The tradeoff is clear: cloud for speed, local for privacy.

The Privacy Argument

This is the part that matters most for a lot of users:

For a journalist working with a confidential source, or a researcher handling IRB-protected interviews, only local mode is acceptable. That's why the tool defaults to Whisper and treats the cloud path as opt-in.

Where This Lives

I packaged this pipeline into a Windows desktop app so non-technical users can run it without touching a terminal:

It's called [Video Transcriber Pro](https://store.rubibot.org/l/videotranscriberpro) and it's a one-time purchase — no subscription, no per-minute billing. The local mode requires no API key at all.

If you just want the code, every piece above is open-source and composable. If you want the polished desktop version, that's the product.

FAQ

*Does local mode really need no internet?

After the first run (which downloads the Whisper model weights, ~1.5 GB for `medium`), yes — fully offline.

*How accurate is Whisper `medium`?

On clean English audio, word error rate is typically 5–8%. Noisy recordings and heavy accents degrade it, but that's true of every ASR system.

*Can I run this on a Mac or Linux?

The pipeline is pure Python and works cross-platform. The desktop app I linked is Windows-only because that's where most of my users are, but the underlying code runs anywhere `ffmpeg` and `whisper` do.

**What about real-time transcription?**

Whisper supports streaming via `whisper-streaming` or the faster-whisper backend. The desktop app currently does batch transcription; real-time is on the roadmap.

Takeaway

Speech-to-text is a solved problem technically. The open question is whether you're willing to upload your audio to solve it. With Whisper running locally, you don't have to — and you still get accuracy good enough for subtitles, meeting notes, and interview transcripts.

The code is open. The packaged app is [product here](https://store.rubibot.org/l/videotranscriberpro) if you want the desktop version.

*If you found this useful, I write about AI tooling and agent architectures at [dev.to/devhunterai](https://dev.to/devhunterai).*
