# How to generate an explainer video with a coding agent and ffmpeg, and what one run costs

> Source: <https://stackness.dev/blog/how-to-generate-an-explainer-video-with-a-coding-agent-and-ffmpeg-and-what-one-run-costs>
> Published: 2026-10-03 10:52:21+00:00

# How to generate an explainer video with a coding agent and ffmpeg, and what one run costs

To generate an explainer video with a coding agent, have the agent write the film as code, then let a headless browser capture it frame by frame and [FFmpeg](https://stackness.dev/tools/ffmpeg) encode the frames. [LaunchVideo](https://stackness.dev/tools/launchvideo), launched on 24 September 2026, does this in about four minutes and roughly 100,000 [Claude Opus](https://stackness.dev/tools/claude-opus) 5.5 tokens per 20 to 40 second film, about $0.71 at list prices by my arithmetic. A [Claude Code](https://stackness.dev/tools/claude-code) run posted to Reddit five days later spent $30.46 in API-equivalent tokens on a 30-second reel. In both, the model writes code and never touches the encoder.

Three projects published numbers that week: LaunchVideo, HN.watch and the Reddit run. For the same length of film, the two code-rendered ones differ in cost by more than 40 times. The Sonnet and Opus pricing math is in the [Sonnet 5.5 vs Opus 5.5 post](https://stackness.dev/blog/sonnet-5-5-vs-opus-5-5-which-slot-each-one-takes-in-a-coding-agent-stack).

## What is in the stack when an agent renders video as code?

Five parts: an agent harness and model, a machine to run on, a headless Chromium with a controlled clock, FFmpeg, and optionally a voice. The agent writes an HTML page, a canvas program or React components. The browser seeks each frame and screenshots it, and FFmpeg encodes the frames to H.264. No video model is involved.

| Layer | LaunchVideo | Reddit reel | HN.watch | 
|---|---|---|---|
| Harness | One [OpenComputer](https://stackness.dev/tools/opencomputer) agent in TypeScript | Claude Code in the Windows desktop app, 14 subagents | Scrimba's own pipeline, or your agent over MCP | 
| Model | Opus 5.5 | [Claude Sonnet](https://stackness.dev/tools/claude-sonnet) 5.5 | "Gemini, GPTs, Inworld, ElevenLabs, and a few others" | 
| Machine | A fresh microVM per job: arm64, 4 vCPU, 8 GB | The author's own PC | Scrimba's servers | 
| Film as code | One HTML file under 60 KB | Canvas scenes, one file per builder agent | HTML slides in Scrimba's player | 
| Capture | [Playwright](https://stackness.dev/tools/playwright) 's headless[Chromium](https://stackness.dev/tools/chromium) , one JPEG per frame | Canvas frames drawn in headless Chrome | None | 
| Encode | FFmpeg, libx264, 1080p at 30 fps | FFmpeg, 1080p at 60 fps | MP4 only on export, renderer not disclosed | 
| Audio | None | A soundtrack synthesised with numpy | TTS narration in 33 languages | 

LaunchVideo's details come from its [repository](https://github.com/diggerhq/shipvideo) and site. The model gets three tools: `web_fetch`, `check_scene` and `render_video`. The Reddit stack is from the author's [r/ClaudeCode write-up](https://www.reddit.com/r/ClaudeCode/comments/1wtapu5/) of 29 September. HN.watch is the outlier. Per Borgen, Scrimba's founder, [wrote on Hacker News](https://news.ycombinator.com/item?id=49879401) that it "doesn't actually capture it as a video, it's just HTML with a voice over essentially".

Both vendor projects carry a conflict of interest. LaunchVideo was posted by a co-founder of OpenComputer, and its footer reads "Built on OpenComputer". HN.watch is a demo of [Scrimba Explain](https://stackness.dev/tools/scrimba-explain). Every number from either is self-reported.

## Which steps does the agent do, and which stay deterministic?

The agent researches, writes the script and the scene code, and checks its own work. Everything that turns code into pixels stays deterministic: a virtual clock, one screenshot per frame, the FFmpeg encode and, in voiced pipelines, scene timing taken from the narration's word timestamps. The model never sees a frame unless a tool hands it one.

LaunchVideo's renderer injects a clock before the page loads, so `requestAnimationFrame`, timers, `Date` and CSS animations advance only when the renderer seeks to the next frame. The prompt bans `Math.random`, CSS transitions and reading the system time, "so renders stay deterministic". Its `check_scene` tool reports JavaScript errors and the visible text at sample timestamps, never an image.

The Reddit run went the other way on review. Five builder agents each wrote one scene, five reviewer agents "rendered contact sheets, read the PNGs, and returned a structured defect list", and four fixers ran only where a reviewer asked. That is Claude Code's [dynamic workflows](https://code.claude.com/docs/en/workflows) feature, run in its ultracode setting.

Voiced pipelines work audio first. [videowright](https://stackness.dev/tools/videowright) generates the narration, takes word timestamps and fits each scene to them. Kiln, which built it, [wrote](https://kiln.tech/blog/we_made_our_launch_video_in_claude_code) that its first launch video "took 2 days of iteration" with the audio added after the picture, and the next one about two hours.

## What did one published run cost, and where did the money go?

Almost all of it went to model tokens. LaunchVideo's roughly 100,000 tokens come to about $0.66 at Opus 5.5 list prices, plus about $0.05 of machine time. The Reddit run's 30-second cut cost $30.46 in API-equivalent Sonnet 5.5 tokens, and 57% of the session's bill was cache reads. Narration, where a pipeline has it, costs cents.

| Run | Published figure | Reported by | Per | 
|---|---|---|---|
| LaunchVideo | About 100k tokens (90k in, 15k out), about 4 minutes. No dollar figure | OpenComputer, on its product page | 20 to 40 second film | 
| LaunchVideo, my arithmetic | About $0.66 in tokens and $0.05 of machine time, $0.71 in all | Derived from list prices | Same | 
| Reddit reel | $4.94 for a 15-second cut, $30.46 for the 30-second cut, $35.40 in all | u/oxmannnn, from transcripts, on a Max 20x plan | Session, two cuts | 
| HN.watch | "~$0.04" in the post, "~$0.4" in a reply the same day | Per Borgen, Scrimba | Video, excluding images | 

The machine line uses [OpenComputer's](https://opencomputer.dev) rate of $0.0126 a minute for 8 GB and 4 vCPU. LaunchVideo does not publish its cache use, so the $0.66 assumes none.

The two HN.watch figures are 10x apart, both from Borgen, and neither is broken down. He wrote that Scrimba does not use Opus 5.5 because "it would be too expensive", and that the goal is "1 cent for a 1 minute video".

Voice is the cheap line. On 3 October, [ElevenLabs](https://stackness.dev/tools/elevenlabs) [charges](https://elevenlabs.io/pricing/api) $0.08 per 1,000 characters on its v3 model, about one minute of speech by its own table, and [Gemini 3.8 Flash TTS](https://ai.google.dev/gemini-api/docs/pricing) works out near $0.014 a minute. In one [r/ClaudeAI run](https://www.reddit.com/r/ClaudeAI/comments/1wogab3/) that broke it out, TTS was $0.23 of roughly $23.

Most of the 40x gap between LaunchVideo and the Reddit reel is the review loop. LaunchVideo writes one file and checks text, while the Reddit run had 14 agents reading images and re-reading a growing context. The render itself costs pennies or runs locally.

## Why was Opus 5.5 only 43% more expensive than Sonnet 5.5 for the same work?

Because 57% of the Sonnet bill was cache reads, which cost $0.20 per million tokens on both models. Every other token class costs twice as much on Opus 5.5, so the same tokens cost 1.43 times as much on Opus, not 2 times. The Opus figure of $50.48 is the author's calculation on Sonnet's tokens, not a second run.

The general rule, worked through in the Sonnet vs Opus post: with cache reads at share *s* of the Sonnet bill, Opus costs 2 minus *s* times as much. Cache reads pile up in a multi-agent render because every API call re-reads the conversation so far, and the 30-second cut made 539 calls. On subscription limits the gap was wider: the author said an earlier Opus 5.5 run of the same video "took ~5% of weekly limit and sonnet 5.5 around 2%".

## How long does one run take, and what is the agent waiting on?

LaunchVideo takes about four minutes: about a minute installing Chromium, FFmpeg and fonts, 30 to 40 seconds to render a 30-second film, and the rest writing roughly 15,000 tokens of HTML and checking it. The Reddit session took about 2.4 hours for both cuts, and the author wrote that "much of that was waiting on renders".

The Reddit render was slow because it was heavy. 1,800 frames at 60 fps with motion blur took about 6.5 minutes, roughly 13 times real time. The author ran renders as background jobs watched by Claude Code's Monitor tool, because "the harness blocks foreground sleep". The 30-second cut alone took about 45 minutes with 14 agents.

HN.watch quotes "a few seconds from click to playback", because it renders no file until someone asks for an MP4.

## Where does this workflow break today?

At the parts the model cannot see. LaunchVideo films are silent and its checks read text, not pixels. Reviewers that read frames catch more and cost more. Pacing, visual quality and one-frame crash bugs are the common failures, along with the bill when a run has to be repeated.

- **No sound.** LaunchVideo's tools throw on`<audio>` ,`<video>` or`<iframe>` . "I expected sound, with an explainer video", one commenter in the[HN thread](https://news.ycombinator.com/item?id=49836374) wrote.
- **Wrong subject.** One user pointed LaunchVideo at a board game reference site and got "a video about an AI startup".
- **Pacing.** The Reddit prompt asked for "not too fast". "I am not sure how you got this from 'not too fast'", a commenter replied.
- **One-frame crashes.** A reviewer in the Reddit run found an`arc()` call with a negative radius that "would have killed the final render at one exact frame".
- **Visuals against the narration.** On HN.watch, one user found "graphics and charts that confuse more than help, or even contradict what is being said". Borgen agreed: "We are not particularly pleased with the state of our visualizations."
- **Reproducibility.** Readers who reused the Reddit prompt reported worse output, and one ran it "on ultracode mode for 7 hours". The author also said he once had to stop the agent pulling assets from another folder.
- **Framework fit.** One HN commenter found Claude models "extremely bad at manim", the[Manim](https://stackness.dev/tools/manim) Python animation library.

## Which open-source projects package this pattern?

| Project | Renders with | Voice | Licence | 
|---|---|---|---|
| [Remotion](https://stackness.dev/tools/remotion) | React in headless Chrome, FFmpeg | Captions skill | Source-available, paid company licence above 3 employees | 
| [HyperFrames](https://stackness.dev/tools/hyperframes) | HTML and GSAP, [Puppeteer](https://stackness.dev/tools/puppeteer) , FFmpeg | TTS and captions skill | Apache 2.0 | 
| [OpenMontage](https://stackness.dev/tools/openmontage) | Remotion or HyperFrames, FFmpeg | Piper by default, ElevenLabs optional | AGPL 3.0 | 
| videowright | Web components, Playwright, FFmpeg | ElevenLabs with word timestamps | MIT | 
| [Motion Canvas](https://stackness.dev/tools/motion-canvas) | TypeScript on canvas | None built in | MIT | 
| Manim | Python scenes, FFmpeg | None built in | MIT | 

OpenMontage is the only one that publishes a cost per video: estimates of $0 to about $2.50 per prompt in its gallery, and a default budget cap of $10. Its [issue #20](https://github.com/calesthio/OpenMontage/issues/20) notes that the estimates leave out failed generations, which providers still bill.

On Stackness, as of 3 October 2026, FFmpeg is on 1 real profile and Playwright, Puppeteer and Manim are on none, against 8 for Claude Code ([data sources](https://stackness.dev/about/data-sources)). The numbers are small, but people list the agent and not the render half of this stack. The [AI tools developers list on Stackness](https://stackness.dev/categories/ai-tools) will show if that changes.

## Key numbers

- **About 100k tokens** and**about 4 minutes** per LaunchVideo film on Opus 5.5, self-reported by OpenComputer (launchvideo.io,**24 September 2026** ), about $0.71 at list prices by my arithmetic.

- **57%** of that bill was cache reads, at**$0.20** per million on both Sonnet 5.5 and Opus 5.5, so the same tokens cost**43%** more on Opus.

- **6.5 minutes** to render 1,800 frames at 60 fps in the Reddit run, against**30 to 40 seconds** for a 30-second LaunchVideo film at 30 fps.
- **1** real Stackness profile lists FFmpeg and**0** list Playwright, Puppeteer or Manim, as of**3 October 2026** ([data sources](https://stackness.dev/about/data-sources) ).

## Quick answers

**How do you generate an explainer video with a coding agent and ffmpeg?** Have the agent write the film as an HTML page, a canvas program or a Remotion project, capture it frame by frame in headless Chromium with a controlled clock, and encode the frames with FFmpeg. LaunchVideo and videowright package the whole loop.

**How much does an agent-made explainer video cost?** From under a dollar to tens of dollars in model tokens. LaunchVideo uses about 100,000 Opus 5.5 tokens per film, about $0.71, and a 14-agent Sonnet 5.5 run spent $30.46 on 30 seconds.

**Does the model render the video?** No. The model writes code, and a headless browser and FFmpeg turn it into frames and an MP4. The model only sees a frame if a tool shows it one.

**Why are cache reads most of the bill?** Every API call re-reads the conversation so far from the prompt cache. In a run with many agents and hundreds of calls, those reads outgrow the output tokens.

**Can a coding agent add narration?** Yes, through a TTS API such as ElevenLabs or Gemini TTS, at a few cents per minute of speech. LaunchVideo films are silent. videowright and OpenMontage generate the voice first and time the scenes to it.

## Tools in this post

## Use any of these tools?

Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.

[Show my stack](https://stackness.dev/register)
