cd /news/ai-tools/scraping-youtube-transcripts-at-scal… · home › topics › ai-tools › article
[ARTICLE · art-140841] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Scraping YouTube transcripts at scale for RAG (and the 3 things that break it)

A developer outlined a Python pipeline for scraping YouTube captions at scale to feed RAG and agent systems, using the public player response and timedtext resources rather than the YouTube Data API v3, whose captions.download endpoint only works for videos owned by the authenticated channel. The approach calls an Apify Actor to handle player payloads, retries and proxy rotation, then chunks timestamped segments into roughly 1,200-character records for embedding. The writeup flags three failure modes, including videos with no captions at all, which require a separate ASR step rather than silently returning an empty document.

by read5 min views1 publishedSep 28, 2026

If you are building a RAG or agent pipeline over video, the transcript is the payload. Titles and descriptions are thin; a 40-minute talk is thousands of tokens of dense, quotable prose. The good news is that YouTube already exposes that prose as captions. The bad news is that pulling captions for videos you do not own, at batch scale, is where most naive pipelines fall over. This post walks through the source-of-truth question, the official-API dead end, a working Python path into a vector store, and the failure modes worth designing around.

You can run speech-to-text on the audio. For a single video it is fine. For hundreds it is expensive, slow, and introduces a second source of error: your ASR output will disagree with the captions every viewer already sees. YouTube captions are aligned to the media, carry per-segment timestamps, and are often human-authored for the videos that matter most. When a human track is missing, auto-generated captions are still a far better starting point than re-transcribing the audio yourself.

One caveat to state plainly: caption extraction is not transcription. If a video has no captions at all, there is nothing to extract — you would need a real ASR step, and you should handle that case explicitly rather than silently returning an empty document.

The obvious first move is the YouTube Data API v3. Its captions.download endpoint looks promising, but it only returns caption tracks for videos owned by the authenticated channel, or where you have permission. For arbitrary third-party videos — the ones your RAG corpus actually needs — it returns an error. The Data API is a channel-management API, not a transcript API.

So the practical route is the public player surface: the watch page's player response carries the caption track metadata, which points at timedtext resources you can fetch. That is the approach YouTube's own web player uses, and it needs no OAuth and no Google Cloud key.

Rather than maintain your own signed HTTP calls, you can call an Actor that already handles the player payload, retries, and proxy rotation. Here is the whole loop: submit a batch, read the dataset, and turn each transcript into chunk records ready for embedding.

from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")

run = client.actor("casaucao/youtube-transcript-scraper").call(run_input={
    "videoUrls": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://youtu.be/9bZkp7q19f0",
    ],
    "languages": ["en", "es"],
    "outputFormat": "segments",
    "includeMetadata": True,
    "maxRetries": 3,
    "concurrency": 5,
})

chunks = []
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    if item.get("status") != "ok" or not item.get("fullText"):
        continue
    segments = item["segments"]
    window = []
    size = 0
    start = segments[0]["start"] if segments else None
    for seg in segments:
        window.append(seg["text"])
        size += len(seg["text"])
        if size >= 1200:
            chunks.append({
                "videoId": item["videoId"],
                "title": item["title"],
                "language": item["language"],
                "start": start,
                "text": " ".join(window),
            })
            window, size = [], 0
            start = seg["end"]
    if window:
        chunks.append({
            "videoId": item["videoId"],
            "title": item["title"],
            "language": item["language"],
            "start": start,
            "text": " ".join(window),
        })

print(len(chunks), "chunks ready to embed")

The fields that matter here: fullText for the whole document, segments for timestamp-preserving chunking, language and translatedFrom for provenance, and isGenerated to know whether a human or the machine wrote it. Once you have chunks, hand each text to your embedding model and store the videoId and title as metadata so citations can link back to the source second.

1. IP blocking at volume. Caption endpoints start rejecting or throttling requests once you push real batch sizes from a single address. Retries help, but only if you rotate the exit IP. Run behind a proxy pool and switch to residential if you are consistently getting block pages. Exponential backoff with jitter keeps a single bad window from cascading into a failed run.

2. The player payload changes underneath you. This is the fragile part of the approach, and it is worth knowing the specific scars. Using the watch-page payload rather than the InnerTube /player endpoint avoids a class of auth problems, but client versions go stale — an outdated InnerTube client version can make playlist browse calls return HTTP 500, and resolving a channel @handle needs a consent cookie in the request before it will resolve. These are not documented behaviors; they are found in production and fixed in production.

3. Per-item failure must not kill the run. With a batch of hundreds of URLs, some will be age-restricted, region-locked, private, or simply caption-less. A pipeline that raises on the first bad item is useless. Each row should carry its own status and error so the caller can log the failures and keep the good transcripts. Translation fallback belongs in the same bucket: if your preferred language is unavailable, fall back to another track and translate when possible, recording translatedFrom so you know the text is not in the original tongue.

Batch extraction is metered per event rather than per compute minute. As of writing the Actor bills $3 per 1,000 transcripts plus $0.50 per 1,000 metadata-enriched videos, and failed videos or videos without captions are not charged. On the free plan, runs are capped to the first few videos so you can validate shape and quality before scaling up. Do that: run three videos, check that segments align with what you hear, confirm the language fields, and only then point it at your whole channel.

If you want to skip the plumbing, I built the Actor used above — the YouTube Transcript Scraper on Apify. It handles single videos, batches, Shorts, playlists and channel expansion, with segments, text, markdown, srt and vtt output. It is my own tool, so treat this as disclosure rather than an impartial review — but the failure modes above are the real ones I hit building it, and they apply to any caption pipeline you write yourself.

── more in #ai-tools 4 stories · sorted by recency
── more on @youtube 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/scraping-youtube-tra…] indexed:0 read:5min 2026-09-28 · —