How to transcribe audio and video files to text with one API call (MP3, MP4, Google Drive, Dropbox) A developer built Audio & Video to Text, an Apify actor that transcribes MP3, MP4, Google Drive and Dropbox files in a single API call, handling audio extraction, chunking, retries and timestamps. The tool uses Whisper large-v3 and is priced at $0.003 per audio minute, with failed files not charged; the developer reports a 50-minute Dropbox MP3 transcribed in about 40 seconds for $0.15. Speech-to-text models are very good now. The annoying part is everything around them: recordings that are too big to upload, video files, and files that live behind a Google Drive or Dropbox share link. This post covers the do-it-yourself route first, then a one-call alternative I built for myself. If you have a short MP3 and Python, open-source Whisper is enough: pip install -U openai-whisper whisper interview.mp3 --model small --output format txt Whisper wants audio. Pull the audio track out with ffmpeg first, and shrink it while you are there, since speech models work on 16 kHz mono anyway: ffmpeg -i meeting.mp4 -vn -ac 1 -ar 16000 -b:a 48k meeting.mp3 A one-hour video becomes an MP3 of about 20 MB. Hosted speech-to-text APIs are much faster than a laptop, but they cap the upload size. OpenAI's transcription endpoint, for example, accepts files up to 25 MB. For a long recording you have to: ffmpeg can do the split: ffmpeg -i meeting.mp3 -f segment -segment time 600 -c copy chunk %03d.mp3 The rest is glue code, plus retries for when the API rate-limits you halfway through. A Google Drive or Dropbox share link opens a preview page, not the file. You need the direct-download form: dl=0 to dl=1 at the end of the link. https://drive.usercontent.google.com/download?id=FILE ID&export=download&confirm=t . None of this is hard. It is just a lot of small steps to maintain. I packaged those steps into a tool: Audio & Video to Text https://apify.com/spokentext/audio-video-to-text on Apify. You give it file links, Drive or Dropbox share links, or an uploaded file, and it handles extraction, splitting, retries and timestamps. Transcription uses Whisper large-v3. curl -X POST "https://api.apify.com/v2/acts/spokentext~audio-video-to-text/run-sync-get-dataset-items?token=YOUR TOKEN" \ -H "Content-Type: application/json" \ -d '{ "urls": "https://example.com/interview.mp3" , "includeSrt": true }' The response is one JSON object per file: { "status": "ok", "fileName": "interview.mp3", "language": "English", "durationSeconds": 1842, "transcribedMinutes": 31, "text": "Thanks for joining me today. Let's start with...", "segments": { "start": 0, "end": 3.4, "text": "Thanks for joining me today." } , "srt": "1\n00:00:00,000 -- 00:00:03,400\nThanks for joining me today.\n" } To process a batch from Python: python from apify client import ApifyClient client = ApifyClient "YOUR APIFY TOKEN" run = client.actor "spokentext/audio-video-to-text" .call run input={ "urls": "https://example.com/interview.mp3", "https://www.dropbox.com/scl/fi/abc123/meeting.mp4?rlkey=xyz&dl=0", , } for item in client.dataset run.default dataset id .iterate items : print item "fileName" , item "status" , len item.get "text", "" From my own runs: a 50-minute MP3 shared through Dropbox was transcribed in about 40 seconds and cost $0.15. Pricing is $0.003 per audio minute, and files that fail are not charged. | Situation | Route | |---|---| | A few short files, privacy matters most | Whisper on your own machine | | Long recordings, video, share links, or batches | The API | | You need speaker labels | Neither; look for a tool with diarization | Disclosure: I built the tool described in "The one-call route". This article was drafted with AI assistance and checked by me.