# How to transcribe audio and video files to text with one API call (MP3, MP4, Google Drive, Dropbox)

> Source: <https://dev.to/clem616/how-to-transcribe-audio-and-video-files-to-text-with-one-api-call-mp3-mp4-google-drive-dropbox-4dml>
> Published: 2026-10-05 16:36:33+00:00

Speech-to-text models are very good now. The annoying part is everything around them: recordings that are too big to upload, video files, and files that live behind a Google Drive or Dropbox share link.

This post covers the do-it-yourself route first, then a one-call alternative I built for myself.

If you have a short MP3 and Python, open-source Whisper is enough:

```
pip install -U openai-whisper
whisper interview.mp3 --model small --output_format txt
```

Whisper wants audio. Pull the audio track out with ffmpeg first, and shrink it while you are there, since speech models work on 16 kHz mono anyway:

```
ffmpeg -i meeting.mp4 -vn -ac 1 -ar 16000 -b:a 48k meeting.mp3
```

A one-hour video becomes an MP3 of about 20 MB.

Hosted speech-to-text APIs are much faster than a laptop, but they cap the upload size. OpenAI's transcription endpoint, for example, accepts files up to 25 MB. For a long recording you have to:

ffmpeg can do the split:

```
ffmpeg -i meeting.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3
```

The rest is glue code, plus retries for when the API rate-limits you halfway through.

A Google Drive or Dropbox share link opens a preview page, not the file. You need the direct-download form:

`dl=0` to `dl=1` at the end of the link.`https://drive.usercontent.google.com/download?id=FILE_ID&export=download&confirm=t`.
None of this is hard. It is just a lot of small steps to maintain.

I packaged those steps into a tool: [Audio & Video to Text](https://apify.com/spokentext/audio-video-to-text) on Apify. You give it file links, Drive or Dropbox share links, or an uploaded file, and it handles extraction, splitting, retries and timestamps. Transcription uses Whisper large-v3.

```
curl -X POST "https://api.apify.com/v2/acts/spokentext~audio-video-to-text/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "urls": ["https://example.com/interview.mp3"], "includeSrt": true }'
```

The response is one JSON object per file:

```
{
    "status": "ok",
    "fileName": "interview.mp3",
    "language": "English",
    "durationSeconds": 1842,
    "transcribedMinutes": 31,
    "text": "Thanks for joining me today. Let's start with...",
    "segments": [{ "start": 0, "end": 3.4, "text": "Thanks for joining me today." }],
    "srt": "1\n00:00:00,000 --> 00:00:03,400\nThanks for joining me today.\n"
}
```

To process a batch from Python:

``` python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("spokentext/audio-video-to-text").call(run_input={
    "urls": [
        "https://example.com/interview.mp3",
        "https://www.dropbox.com/scl/fi/abc123/meeting.mp4?rlkey=xyz&dl=0",
    ],
})

for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item["fileName"], item["status"], len(item.get("text", "")))
```

From my own runs: a 50-minute MP3 shared through Dropbox was transcribed in about 40 seconds and cost $0.15. Pricing is $0.003 per audio minute, and files that fail are not charged.

| Situation | Route | 
|---|---|
| A few short files, privacy matters most | Whisper on your own machine | 
| Long recordings, video, share links, or batches | The API | 
| You need speaker labels | Neither; look for a tool with diarization | 

*Disclosure: I built the tool described in "The one-call route". This article was drafted with AI assistance and checked by me.*
