How to Make Videos with Claude Code Developer Vincent Schumacher released explainroo, a free, MIT-licensed open-source framework that lets Claude Code and other AI coding agents such as Codex, Pi and OpenCode produce animated explainer videos by writing HTML/JavaScript canvas scenes that headless Chrome renders frame by frame and ffmpeg encodes into MP4. The tool runs its voice model, word timing and music creation locally with no API keys, requires Node.js 20.11 or newer, ffmpeg and Chrome or Chromium, and downloads about 400 MB of voice and speech models via the `doctor --fetch` command; a 15-second video renders in 16 seconds, and no graphics card is needed. OpenAI Is Paying Me to Move My Business to Anthropic Three months ago, OpenAI was my daily driver, using Codex with GPT-5.5 and being productive like never before. Anthropic already had Fable… Claude Code can't generate video directly. It isn't a video model like Veo or Sora. It writes text and code. You can still make videos with it, because a video is a series of images plus an audio track, and code can draw images. The usual way to do it is HTML and JavaScript. A web page with a canvas draws the picture for any point in time. A headless Chrome steps through the video frame by frame and saves each picture, and ffmpeg turns the pictures into an MP4. Claude writes the page, and a script does the rendering. I expected this to look like a slide deck. But actually the videos look much better than that. Lines draw themselves, icons and boxes animate in, the camera moves and zooms, and text can appear word by word as the voice says it. With a voice, music and sound effects on top, you get a proper animated video. To make the whole process of creating videos with Claude Code or any AI Coding Agent, so Codex, Pi, OpenCode etc. work, too I built a framework for it, explainroo. It is free and open source, and the voice model, the word timing and the music creation run on your computer, so it needs no API keys. I made this video with explainroo. The code is on GitHub https://github.com/vincentsch/explainroo under the MIT license, and there are more example videos on explainroo.com https://www.explainroo.com/videos/ . Open Claude Code, paste this and put your topic in place of the brackets: Make me a short explainer video about your topic . Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps. The agent clones the repo, sets it up, writes the video, checks it and renders an MP4. Codex and other coding agents work too. If you want to set it up yourself, you need Node.js 20.11 or newer, ffmpeg, and Chrome or Chromium: git clone https://github.com/vincentsch/explainroo.git cd explainroo npm install node bin/explainroo.js doctor --fetch doctor --fetch downloads the voice model and the speech model once, about 400 MB together. You don't need a graphics card. I develop and test on Linux. macOS and Windows should work, but I have tested them less. A video is a folder with two files. script.md holds the narration, one section per scene: hook Most people never think about what happens click after they click a link. steps It takes three steps. First, one your browser finds the server. Then two it asks for the page. Finally, three it draws what comes back. one is a marker. It is not spoken, and it gets the time of the word that follows it. {SQL|sequel} shows "SQL" in the captions and makes the voice say "sequel", for words the voice reads wrong. scenes.js has one function per scene: export default { hook s { s.title 'What happens after a click?', { at: 0 } ; s.icon 'mouse-pointer-click', { y: 760, size: 130, color: 'accent', at: ' click' } ; }, steps s { s.box 'Find the server', { id: 'find', x: 420, y: 540, icon: 'search', at: ' one' } ; s.box 'Ask for the page', { id: 'ask', x: 960, y: 540, icon: 'send', at: ' two' } ; s.box 'Draw it', { id: 'draw', x: 1500, y: 540, icon: 'paintbrush', at: 'draws' } ; s.arrow 'find', 'ask', { at: ' two' } ; s.arrow 'ask', 'draw', { at: 'draws' } ; }, }; at takes seconds, a marker ' one' or a spoken word 'draws' . This is the last frame of the second scene. The whole 15 second video renders in 16 seconds. The engine calls the scene function once for every frame and passes a fresh s with the scene time s.t . Nothing is kept between frames. Each element works out its own state from the time. Before its at time it draws nothing. After that, its progress is the time since at divided by the length of its animation, passed through an easing curve. After its out time it fades or slides away. The look decides how an element enters. In the hand-drawn looks, shapes "draw" in: the stroke is revealed with setLineDash over the measured path length during the first 78% of the animation, the fill starts at 42% and the label at 35%. Because a frame depends only on its time, any frame can be rendered on its own, in any order and in parallel. That is also why the checks below can jump to any moment. Hand-drawn lines come from Rough.js https://roughjs.com , which adds random wobble to every stroke. If the wobble changed from frame to frame, every line would jitter. So each element gets a seed from its scene and its id, and while a scene draws, Math.random is replaced by a seeded generator. The lines stay still. A boil setting changes the seeds a few times per second if you want the jittery look on purpose. The 1,854 icons are Lucide https://lucide.dev SVG paths on a 24 by 24 grid. The engine scales the path data to the icon size and draws the paths with the same pen as the shapes, so icons draw themselves in like everything else. A scene transition draws both scenes into two offscreen canvases and combines them with a fade, slide, wipe, zoom or brush stroke. Captions and the watermark go on top. The voice is Kokoro https://github.com/hexgrad/kokoro , an open text to speech model with 82 million parameters. It runs in Node through kokoro-js, as an ONNX model in fp32 on the CPU, and outputs 24 kHz audio. It has 28 voices. The narration is cut into sentences, and each sentence is synthesized on its own. Leading and trailing silence is trimmed, and the engine puts 0.3 seconds of silence between sentences and 0.55 seconds between paragraphs. That keeps the pauses the same length everywhere, whatever the model does. Voice files are cached per scene. The cache key is a SHA-256 over the model names, the voice, the speed, the sentences and the positions of the markers. When you change one scene, only that scene is spoken again. Kokoro doesn't say when each word was spoken, so each sentence goes through speech recognition afterwards. It is resampled to 16 kHz and transcribed with whisper-base.en https://huggingface.co/onnx-community/whisper-base.en timestamped , 8-bit quantized, through Transformers.js https://github.com/huggingface/transformers.js , with word timestamps. Then the words of the script are lined up with the words Whisper heard. This is dynamic programming over a cost matrix, like an edit distance. A matching word costs 0, a skipped script word or heard word costs 1, and a different word costs 1.6. Two extra moves handle words that are split differently on both sides, like "D N S" against "DNS". For numbers, both sides are expanded into spoken words before they are compared, because the voice says "one thousand and fifty dollars" while Whisper writes "$1,050", sometimes in pieces. The "and" inside a number is optional. Words that Whisper didn't confirm get times by interpolation between their confirmed neighbours, weighted by their length in characters. A marker takes the start time of the next word. A scene then lasts 0.35 seconds of lead-in, plus the narration, plus 0.7 seconds of hold, snapped to the 30 fps frame grid. The pace setting speeds up the voice and divides every lead-in, hold, pause and animation by the same factor. The music gets faster by the square root of it. The alignment also works as a pronunciation check. A word only counts as confirmed when every part of it was heard. In my Vroni product demo, the check couldn't confirm "Vroni", because Whisper heard "Rony". Rewording the sentence fixed it. The audio is a synthesizer written in plain JavaScript. I didn't use Web Audio nodes for the mix, because Web Audio sums node inputs in an unspecified order, so two renders of the same video can differ in the last bits. The synth renders oscillators and noise with envelopes, filters, panning and a reverb send in blocks of 256 samples at 48 kHz, and it always produces the same samples for the same input. The same code runs in the preview page and in the final render. The music is generated for each video. There are six styles. A style has a tempo, a set of chord progressions and a function that plays one bar on its instruments. A scheduler adds sections, voice leading between chords, small accents on scene changes and a final cadence. A seed picks the key and the variations, so the same settings always give the same music. Kokoro's narration lands at about -20 LUFS. At the default volume the music sits about 12 LU below it, and it ducks another 8 dB while someone speaks, with a 0.2 second attack and a 0.5 second release around each spoken span. For the sound effects, the engine runs each scene at a quarter of the resolution every 0.1 seconds and records which elements enter. Each entrance can trigger a synthesized effect: a pop, a scribble, a click, a key press. Effects are thinned so that at most three play in any half second. A small local HTTP server serves the engine, the project files and the timeline as JSON. Playwright https://playwright.dev starts a headless Chrome and opens one page per worker. There are half as many workers as CPU cores, at most eight. The first page mixes the soundtrack and writes it as a WAV file. ffmpeg normalizes it to -14 LUFS in two runs, one to measure and one to apply the correction, while the frames render. Each page gets its own range of frames. It draws six frames per call, encodes them as JPEG with canvas.toDataURL , and sends them back to Node as base64. Node writes them into a separate ffmpeg process per page: ffmpeg -f image2pipe -framerate 30 -c:v mjpeg -i - \ -c:v libx264 -preset medium -crf 17 -tune animation \ -pix fmt yuv420p -g 60 seg-00.mp4 At the end, ffmpeg's concat demuxer joins the segments with -c:v copy , without encoding them again, and adds the audio as AAC. A --draft render uses half the resolution, CRF 24 and the veryfast preset. On my laptop with 16 threads, the 59 second Unspar product demo renders 1,762 frames at 1080p in 47 seconds. The whole render, including the soundtrack and the final mux, takes 58 seconds. A coding agent can't watch a video, but it can look at images and read text. So there are four commands for checking. still renders PNGs of chosen moments, by default the end of each scene. sheet puts small frames of the whole video into one image, so the agent can see motion. check goes through the video every 0.25 seconds. Each text element logs its bounding box, transformed by the current canvas matrix into frame coordinates. verify checks the finished MP4 with ffmpeg: ebur128 for loudness and true peak, blackdetect for black frames and silencedetect for long silences. It also runs Whisper on the final mix and compares the result with the script, to see if the voice is still clear over the music. verify once reported that only 70% of a video's narration could be understood. The audio was fine. Whisper transcribes long files in 30 second windows, and for one window it returned only " bell dings ". Now verify cuts the mix into scenes using the timeline and transcribes each scene on its own, which also tells you which scene has a problem. For product demos, a UI kit, s.ui , draws app screens: cards, input fields, dropdowns, toggles, buttons and a mouse pointer. They are drawn crisp, in Inter, with serif headlines in Instrument Serif, or in the product's own colors, fonts and logo from video.json . The pointer moves along keyframes with an optional curve and clicks with a sound, text is typed at 17 characters per second with key sounds, and the camera zooms in on keyframes as well. The first demo rendered at 4.9 frames per second. A few blurred app cards drift in the background, and a canvas filter: blur on each frame is very slow. Now each card is blurred once into an OffscreenCanvas and reused. Today the frames of the same preview render in 11 seconds instead of 370. I made demos for two of my own products, Unspar https://www.explainroo.com/videos/unspar-product-demo/ and Vroni https://www.explainroo.com/videos/vroni-product-demo/ . The text on the screens comes from the apps' code, so the labels and buttons match the apps. explainroo.com also has unofficial demos of Gmail, ChatGPT and Claude https://www.explainroo.com/videos/ product-demos , with each feature checked against the official help pages. Give Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request. Take a look at vroni.com