Speech-to-text models are very good now. The annoying part is everything around them: recordings that are too big to upload, video files, and files that live behind a Google Drive or Dropbox share link.
This post covers the do-it-yourself route first, then a one-call alternative I built for myself.
If you have a short MP3 and Python, open-source Whisper is enough:
pip install -U openai-whisper
whisper interview.mp3 --model small --output_format txt
Whisper wants audio. Pull the audio track out with ffmpeg first, and shrink it while you are there, since speech models work on 16 kHz mono anyway:
ffmpeg -i meeting.mp4 -vn -ac 1 -ar 16000 -b:a 48k meeting.mp3
A one-hour video becomes an MP3 of about 20 MB.
Hosted speech-to-text APIs are much faster than a laptop, but they cap the upload size. OpenAI's transcription endpoint, for example, accepts files up to 25 MB. For a long recording you have to:
ffmpeg can do the split:
ffmpeg -i meeting.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3
The rest is glue code, plus retries for when the API rate-limits you halfway through.
A Google Drive or Dropbox share link opens a preview page, not the file. You need the direct-download form:
dl=0 to dl=1 at the end of the link.https://drive.usercontent.google.com/download?id=FILE_ID&export=download&confirm=t.
None of this is hard. It is just a lot of small steps to maintain.
I packaged those steps into a tool: Audio & Video to Text on Apify. You give it file links, Drive or Dropbox share links, or an uploaded file, and it handles extraction, splitting, retries and timestamps. Transcription uses Whisper large-v3.
curl -X POST "https://api.apify.com/v2/acts/spokentext~audio-video-to-text/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "urls": ["https://example.com/interview.mp3"], "includeSrt": true }'
The response is one JSON object per file:
{
"status": "ok",
"fileName": "interview.mp3",
"language": "English",
"durationSeconds": 1842,
"transcribedMinutes": 31,
"text": "Thanks for joining me today. Let's start with...",
"segments": [{ "start": 0, "end": 3.4, "text": "Thanks for joining me today." }],
"srt": "1\n00:00:00,000 --> 00:00:03,400\nThanks for joining me today.\n"
}
To process a batch from Python:
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("spokentext/audio-video-to-text").call(run_input={
"urls": [
"https://example.com/interview.mp3",
"https://www.dropbox.com/scl/fi/abc123/meeting.mp4?rlkey=xyz&dl=0",
],
})
for item in client.dataset(run.default_dataset_id).iterate_items():
print(item["fileName"], item["status"], len(item.get("text", "")))
From my own runs: a 50-minute MP3 shared through Dropbox was transcribed in about 40 seconds and cost $0.15. Pricing is $0.003 per audio minute, and files that fail are not charged.
| Situation | Route |
|---|---|
| A few short files, privacy matters most | Whisper on your own machine |
| Long recordings, video, share links, or batches | The API |
| You need speaker labels | Neither; look for a tool with diarization |
Disclosure: I built the tool described in "The one-call route". This article was drafted with AI assistance and checked by me.