cd /news/ai-tools/how-to-batch-transcribe-audio-files-… · home › topics › ai-tools › article
[ARTICLE · art-141633] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

How to Batch-Transcribe Audio Files on Your Computer With a Local API in 2026

OGAD (Off Grid AI Desktop) version 0.0.51 exposes a local OpenAI-compatible transcription endpoint at /v1/audio/transcriptions, letting users batch-transcribe WAV files on their own machine with a downloaded local model instead of sending audio to an online service. A provided Bash script iterates over an audio folder, writes one .txt transcript per file, and reports failures without leaving partial output; the endpoint caps uploads at 200 MiB and requires no API key, so the port should stay on a trusted network.

by read3 min views1 publishedSep 29, 2026

A folder of recordings should not need a separate manual transcription request for every file.

OGAD (Off Grid AI Desktop) exposes local transcription through /v1/audio/transcriptions. A small script can send recordings one at a time and save a text file for each. With a downloaded local transcription model selected, the audio is processed on your computer instead of an online transcription service.

Download OGAD for Mac or Windows

Install OGAD and download a local transcription model in Models. Select the local model, then open Gateway and check the address. The usual local port is 7878; change the examples if the app uses another one.

For multilingual recordings, select a multilingual transcription model. An English-only model is not the right route for a folder of Hindi or Spanish recordings.

The gateway is part of the core app. You do not need Pro meeting capture to submit existing audio files through this endpoint. Initial model downloads need internet; the configured local transcription route can then work offline.

Try one WAV file first:

curl --fail --show-error \
  http://127.0.0.1:7878/v1/audio/transcriptions \
  -F 'file=@./audio/sample.wav' \
  -F 'response_format=text'

The expected result is the transcript text printed in the terminal. Listen to a short section of the original and check the words before processing the folder.

The following Bash script processes .wav files directly inside an audio folder. It uses a new output directory and refuses to overwrite an existing one.

Save it as transcribe-folder.sh:

#!/usr/bin/env bash
set -u
shopt -s nullglob

base='http://127.0.0.1:7878'
input_dir='./audio'
output_dir='./transcripts'
files=("$input_dir"/*.wav)

if [ "${#files[@]}" -eq 0 ]; then
  echo 'No WAV files found in ./audio' >&2
  exit 1
fi
if ! mkdir "$output_dir"; then
  echo 'Choose a new output directory before running again.' >&2
  exit 1
fi

failed=0
for file in "${files[@]}"; do
  name=$(basename "$file" .wav)
  target="$output_dir/$name.txt"
  temporary="$target.partial"
  echo "Transcribing: $file"
  if curl --fail --show-error --silent \
      "$base/v1/audio/transcriptions" \
      -F "file=@$file" -F 'response_format=text' \
      -o "$temporary"; then
    mv "$temporary" "$target"
  else
    echo "Failed: $file" >&2
    rm -f "$temporary"
    failed=1
  fi
done
exit "$failed"

Run it from the directory that contains audio:

bash transcribe-folder.sh

Each successful request becomes a .txt file. Failed requests are reported and do not leave a completed transcript under the expected name. Keep the original audio until you have checked the results.

The released endpoint limits the complete upload request to 200 MiB, including multipart overhead. Leave room below that limit. Split a larger recording into smaller audio files before submitting it.

This example sends one file at a time. That keeps the batch simple and avoids launching multiple heavy requests merely because the folder has many files. Processing time depends on the selected model, audio length and hardware.

The text response is a transcript. It does not give this script speaker identities, word timestamps or an SRT subtitle file. Add those only through a separately verified route if your workflow needs them.

Names, technical terms and mixed-language speech still need review. Check important passages against the audio before quoting them in a report.

The gateway listens on network interfaces and its inference endpoints do not require an API key. These examples use 127.0.0.1 on the same computer. Keep the host on a trusted network and do not expose this port to the public internet.

These API routes are present in OGAD 0.0.51. The running gateway also serves its API reference at /docs.

Download OGAD, test one recording and run a small batch. Keep a transcript beside each source file so your notes can become searchable without a cloud AI upload.

── more in #ai-tools 4 stories · sorted by recency
── more on @ogad 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-batch-transcr…] indexed:0 read:3min 2026-09-29 · —