How to generate an explainer video with a coding agent and ffmpeg, and what one run costs LaunchVideo, launched 24 September 2026, generates a 20 to 40 second explainer video in about four minutes using roughly 100,000 Claude Opus 5.5 tokens, about $0.71 at list prices, by having a coding agent write the film as HTML and code while Playwright's headless Chromium captures one JPEG per frame and FFmpeg encodes to H.264 at 1080p and 30 fps. A separate Claude Code run posted to Reddit on 29 September spent $30.46 in API-equivalent tokens on a 30-second reel, a more than 40-fold cost difference for the same film length between the two code-rendered projects. LaunchVideo was posted by a co-founder of OpenComputer and its footer reads "Built on OpenComputer," so its figures are self-reported. How to generate an explainer video with a coding agent and ffmpeg, and what one run costs To generate an explainer video with a coding agent, have the agent write the film as code, then let a headless browser capture it frame by frame and FFmpeg https://stackness.dev/tools/ffmpeg encode the frames. LaunchVideo https://stackness.dev/tools/launchvideo , launched on 24 September 2026, does this in about four minutes and roughly 100,000 Claude Opus https://stackness.dev/tools/claude-opus 5.5 tokens per 20 to 40 second film, about $0.71 at list prices by my arithmetic. A Claude Code https://stackness.dev/tools/claude-code run posted to Reddit five days later spent $30.46 in API-equivalent tokens on a 30-second reel. In both, the model writes code and never touches the encoder. Three projects published numbers that week: LaunchVideo, HN.watch and the Reddit run. For the same length of film, the two code-rendered ones differ in cost by more than 40 times. The Sonnet and Opus pricing math is in the Sonnet 5.5 vs Opus 5.5 post https://stackness.dev/blog/sonnet-5-5-vs-opus-5-5-which-slot-each-one-takes-in-a-coding-agent-stack . What is in the stack when an agent renders video as code? Five parts: an agent harness and model, a machine to run on, a headless Chromium with a controlled clock, FFmpeg, and optionally a voice. The agent writes an HTML page, a canvas program or React components. The browser seeks each frame and screenshots it, and FFmpeg encodes the frames to H.264. No video model is involved. | Layer | LaunchVideo | Reddit reel | HN.watch | |---|---|---|---| | Harness | One OpenComputer https://stackness.dev/tools/opencomputer agent in TypeScript | Claude Code in the Windows desktop app, 14 subagents | Scrimba's own pipeline, or your agent over MCP | | Model | Opus 5.5 | Claude Sonnet https://stackness.dev/tools/claude-sonnet 5.5 | "Gemini, GPTs, Inworld, ElevenLabs, and a few others" | | Machine | A fresh microVM per job: arm64, 4 vCPU, 8 GB | The author's own PC | Scrimba's servers | | Film as code | One HTML file under 60 KB | Canvas scenes, one file per builder agent | HTML slides in Scrimba's player | | Capture | Playwright https://stackness.dev/tools/playwright 's headless Chromium https://stackness.dev/tools/chromium , one JPEG per frame | Canvas frames drawn in headless Chrome | None | | Encode | FFmpeg, libx264, 1080p at 30 fps | FFmpeg, 1080p at 60 fps | MP4 only on export, renderer not disclosed | | Audio | None | A soundtrack synthesised with numpy | TTS narration in 33 languages | LaunchVideo's details come from its repository https://github.com/diggerhq/shipvideo and site. The model gets three tools: web fetch , check scene and render video . The Reddit stack is from the author's r/ClaudeCode write-up https://www.reddit.com/r/ClaudeCode/comments/1wtapu5/ of 29 September. HN.watch is the outlier. Per Borgen, Scrimba's founder, wrote on Hacker News https://news.ycombinator.com/item?id=49879401 that it "doesn't actually capture it as a video, it's just HTML with a voice over essentially". Both vendor projects carry a conflict of interest. LaunchVideo was posted by a co-founder of OpenComputer, and its footer reads "Built on OpenComputer". HN.watch is a demo of Scrimba Explain https://stackness.dev/tools/scrimba-explain . Every number from either is self-reported. Which steps does the agent do, and which stay deterministic? The agent researches, writes the script and the scene code, and checks its own work. Everything that turns code into pixels stays deterministic: a virtual clock, one screenshot per frame, the FFmpeg encode and, in voiced pipelines, scene timing taken from the narration's word timestamps. The model never sees a frame unless a tool hands it one. LaunchVideo's renderer injects a clock before the page loads, so requestAnimationFrame , timers, Date and CSS animations advance only when the renderer seeks to the next frame. The prompt bans Math.random , CSS transitions and reading the system time, "so renders stay deterministic". Its check scene tool reports JavaScript errors and the visible text at sample timestamps, never an image. The Reddit run went the other way on review. Five builder agents each wrote one scene, five reviewer agents "rendered contact sheets, read the PNGs, and returned a structured defect list", and four fixers ran only where a reviewer asked. That is Claude Code's dynamic workflows https://code.claude.com/docs/en/workflows feature, run in its ultracode setting. Voiced pipelines work audio first. videowright https://stackness.dev/tools/videowright generates the narration, takes word timestamps and fits each scene to them. Kiln, which built it, wrote https://kiln.tech/blog/we made our launch video in claude code that its first launch video "took 2 days of iteration" with the audio added after the picture, and the next one about two hours. What did one published run cost, and where did the money go? Almost all of it went to model tokens. LaunchVideo's roughly 100,000 tokens come to about $0.66 at Opus 5.5 list prices, plus about $0.05 of machine time. The Reddit run's 30-second cut cost $30.46 in API-equivalent Sonnet 5.5 tokens, and 57% of the session's bill was cache reads. Narration, where a pipeline has it, costs cents. | Run | Published figure | Reported by | Per | |---|---|---|---| | LaunchVideo | About 100k tokens 90k in, 15k out , about 4 minutes. No dollar figure | OpenComputer, on its product page | 20 to 40 second film | | LaunchVideo, my arithmetic | About $0.66 in tokens and $0.05 of machine time, $0.71 in all | Derived from list prices | Same | | Reddit reel | $4.94 for a 15-second cut, $30.46 for the 30-second cut, $35.40 in all | u/oxmannnn, from transcripts, on a Max 20x plan | Session, two cuts | | HN.watch | "~$0.04" in the post, "~$0.4" in a reply the same day | Per Borgen, Scrimba | Video, excluding images | The machine line uses OpenComputer's https://opencomputer.dev rate of $0.0126 a minute for 8 GB and 4 vCPU. LaunchVideo does not publish its cache use, so the $0.66 assumes none. The two HN.watch figures are 10x apart, both from Borgen, and neither is broken down. He wrote that Scrimba does not use Opus 5.5 because "it would be too expensive", and that the goal is "1 cent for a 1 minute video". Voice is the cheap line. On 3 October, ElevenLabs https://stackness.dev/tools/elevenlabs charges https://elevenlabs.io/pricing/api $0.08 per 1,000 characters on its v3 model, about one minute of speech by its own table, and Gemini 3.8 Flash TTS https://ai.google.dev/gemini-api/docs/pricing works out near $0.014 a minute. In one r/ClaudeAI run https://www.reddit.com/r/ClaudeAI/comments/1wogab3/ that broke it out, TTS was $0.23 of roughly $23. Most of the 40x gap between LaunchVideo and the Reddit reel is the review loop. LaunchVideo writes one file and checks text, while the Reddit run had 14 agents reading images and re-reading a growing context. The render itself costs pennies or runs locally. Why was Opus 5.5 only 43% more expensive than Sonnet 5.5 for the same work? Because 57% of the Sonnet bill was cache reads, which cost $0.20 per million tokens on both models. Every other token class costs twice as much on Opus 5.5, so the same tokens cost 1.43 times as much on Opus, not 2 times. The Opus figure of $50.48 is the author's calculation on Sonnet's tokens, not a second run. The general rule, worked through in the Sonnet vs Opus post: with cache reads at share s of the Sonnet bill, Opus costs 2 minus s times as much. Cache reads pile up in a multi-agent render because every API call re-reads the conversation so far, and the 30-second cut made 539 calls. On subscription limits the gap was wider: the author said an earlier Opus 5.5 run of the same video "took ~5% of weekly limit and sonnet 5.5 around 2%". How long does one run take, and what is the agent waiting on? LaunchVideo takes about four minutes: about a minute installing Chromium, FFmpeg and fonts, 30 to 40 seconds to render a 30-second film, and the rest writing roughly 15,000 tokens of HTML and checking it. The Reddit session took about 2.4 hours for both cuts, and the author wrote that "much of that was waiting on renders". The Reddit render was slow because it was heavy. 1,800 frames at 60 fps with motion blur took about 6.5 minutes, roughly 13 times real time. The author ran renders as background jobs watched by Claude Code's Monitor tool, because "the harness blocks foreground sleep". The 30-second cut alone took about 45 minutes with 14 agents. HN.watch quotes "a few seconds from click to playback", because it renders no file until someone asks for an MP4. Where does this workflow break today? At the parts the model cannot see. LaunchVideo films are silent and its checks read text, not pixels. Reviewers that read frames catch more and cost more. Pacing, visual quality and one-frame crash bugs are the common failures, along with the bill when a run has to be repeated. - No sound. LaunchVideo's tools throw on