Scraping YouTube transcripts at scale for RAG (and the 3 things that break it) A developer outlined a Python pipeline for scraping YouTube captions at scale to feed RAG and agent systems, using the public player response and timedtext resources rather than the YouTube Data API v3, whose captions.download endpoint only works for videos owned by the authenticated channel. The approach calls an Apify Actor to handle player payloads, retries and proxy rotation, then chunks timestamped segments into roughly 1,200-character records for embedding. The writeup flags three failure modes, including videos with no captions at all, which require a separate ASR step rather than silently returning an empty document. If you are building a RAG or agent pipeline over video, the transcript is the payload. Titles and descriptions are thin; a 40-minute talk is thousands of tokens of dense, quotable prose. The good news is that YouTube already exposes that prose as captions. The bad news is that pulling captions for videos you do not own, at batch scale, is where most naive pipelines fall over. This post walks through the source-of-truth question, the official-API dead end, a working Python path into a vector store, and the failure modes worth designing around. You can run speech-to-text on the audio. For a single video it is fine. For hundreds it is expensive, slow, and introduces a second source of error: your ASR output will disagree with the captions every viewer already sees. YouTube captions are aligned to the media, carry per-segment timestamps, and are often human-authored for the videos that matter most. When a human track is missing, auto-generated captions are still a far better starting point than re-transcribing the audio yourself. One caveat to state plainly: caption extraction is not transcription. If a video has no captions at all, there is nothing to extract — you would need a real ASR step, and you should handle that case explicitly rather than silently returning an empty document. The obvious first move is the YouTube Data API v3. Its captions.download endpoint looks promising, but it only returns caption tracks for videos owned by the authenticated channel, or where you have permission. For arbitrary third-party videos — the ones your RAG corpus actually needs — it returns an error. The Data API is a channel-management API, not a transcript API. So the practical route is the public player surface: the watch page's player response carries the caption track metadata, which points at timedtext resources you can fetch. That is the approach YouTube's own web player uses, and it needs no OAuth and no Google Cloud key. Rather than maintain your own signed HTTP calls, you can call an Actor that already handles the player payload, retries, and proxy rotation. Here is the whole loop: submit a batch, read the dataset, and turn each transcript into chunk records ready for embedding. python from apify client import ApifyClient client = ApifyClient "