{"slug": "how-to-transcribe-audio-and-video-files-to-text-with-one-api-call-mp3-mp4-google", "title": "How to transcribe audio and video files to text with one API call (MP3, MP4, Google Drive, Dropbox)", "summary": "A developer built Audio & Video to Text, an Apify actor that transcribes MP3, MP4, Google Drive and Dropbox files in a single API call, handling audio extraction, chunking, retries and timestamps. The tool uses Whisper large-v3 and is priced at $0.003 per audio minute, with failed files not charged; the developer reports a 50-minute Dropbox MP3 transcribed in about 40 seconds for $0.15.", "body_md": "Speech-to-text models are very good now. The annoying part is everything around them: recordings that are too big to upload, video files, and files that live behind a Google Drive or Dropbox share link.\n\nThis post covers the do-it-yourself route first, then a one-call alternative I built for myself.\n\nIf you have a short MP3 and Python, open-source Whisper is enough:\n\n```\npip install -U openai-whisper\nwhisper interview.mp3 --model small --output_format txt\n```\n\nWhisper wants audio. Pull the audio track out with ffmpeg first, and shrink it while you are there, since speech models work on 16 kHz mono anyway:\n\n```\nffmpeg -i meeting.mp4 -vn -ac 1 -ar 16000 -b:a 48k meeting.mp3\n```\n\nA one-hour video becomes an MP3 of about 20 MB.\n\nHosted speech-to-text APIs are much faster than a laptop, but they cap the upload size. OpenAI's transcription endpoint, for example, accepts files up to 25 MB. For a long recording you have to:\n\nffmpeg can do the split:\n\n```\nffmpeg -i meeting.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3\n```\n\nThe rest is glue code, plus retries for when the API rate-limits you halfway through.\n\nA Google Drive or Dropbox share link opens a preview page, not the file. You need the direct-download form:\n\n`dl=0` to `dl=1` at the end of the link.`https://drive.usercontent.google.com/download?id=FILE_ID&export=download&confirm=t`.\nNone of this is hard. It is just a lot of small steps to maintain.\n\nI packaged those steps into a tool: [Audio & Video to Text](https://apify.com/spokentext/audio-video-to-text) on Apify. You give it file links, Drive or Dropbox share links, or an uploaded file, and it handles extraction, splitting, retries and timestamps. Transcription uses Whisper large-v3.\n\n```\ncurl -X POST \"https://api.apify.com/v2/acts/spokentext~audio-video-to-text/run-sync-get-dataset-items?token=YOUR_TOKEN\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{ \"urls\": [\"https://example.com/interview.mp3\"], \"includeSrt\": true }'\n```\n\nThe response is one JSON object per file:\n\n```\n{\n    \"status\": \"ok\",\n    \"fileName\": \"interview.mp3\",\n    \"language\": \"English\",\n    \"durationSeconds\": 1842,\n    \"transcribedMinutes\": 31,\n    \"text\": \"Thanks for joining me today. Let's start with...\",\n    \"segments\": [{ \"start\": 0, \"end\": 3.4, \"text\": \"Thanks for joining me today.\" }],\n    \"srt\": \"1\\n00:00:00,000 --> 00:00:03,400\\nThanks for joining me today.\\n\"\n}\n```\n\nTo process a batch from Python:\n\n``` python\nfrom apify_client import ApifyClient\n\nclient = ApifyClient(\"YOUR_APIFY_TOKEN\")\n\nrun = client.actor(\"spokentext/audio-video-to-text\").call(run_input={\n    \"urls\": [\n        \"https://example.com/interview.mp3\",\n        \"https://www.dropbox.com/scl/fi/abc123/meeting.mp4?rlkey=xyz&dl=0\",\n    ],\n})\n\nfor item in client.dataset(run.default_dataset_id).iterate_items():\n    print(item[\"fileName\"], item[\"status\"], len(item.get(\"text\", \"\")))\n```\n\nFrom my own runs: a 50-minute MP3 shared through Dropbox was transcribed in about 40 seconds and cost $0.15. Pricing is $0.003 per audio minute, and files that fail are not charged.\n\n| Situation | Route | \n|---|---|\n| A few short files, privacy matters most | Whisper on your own machine | \n| Long recordings, video, share links, or batches | The API | \n| You need speaker labels | Neither; look for a tool with diarization | \n\n*Disclosure: I built the tool described in \"The one-call route\". This article was drafted with AI assistance and checked by me.*", "url": "https://wpnews.pro/news/how-to-transcribe-audio-and-video-files-to-text-with-one-api-call-mp3-mp4-google", "canonical_source": "https://dev.to/clem616/how-to-transcribe-audio-and-video-files-to-text-with-one-api-call-mp3-mp4-google-drive-dropbox-4dml", "published_at": "2026-10-05 16:36:33+00:00", "updated_at": "2026-10-05 16:48:11.651723+00:00", "lang": "en", "topics": ["ai-tools", "natural-language-processing", "developer-tools"], "entities": ["Apify", "Whisper", "OpenAI", "Google Drive", "Dropbox", "ffmpeg"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-to-transcribe-audio-and-video-files-to-text-with-one-api-call-mp3-mp4-google", "markdown": "https://wpnews.pro/news/how-to-transcribe-audio-and-video-files-to-text-with-one-api-call-mp3-mp4-google.md", "text": "https://wpnews.pro/news/how-to-transcribe-audio-and-video-files-to-text-with-one-api-call-mp3-mp4-google.txt", "jsonld": "https://wpnews.pro/news/how-to-transcribe-audio-and-video-files-to-text-with-one-api-call-mp3-mp4-google.jsonld"}}