cd /news/ai-tools/how-to-render-a-motion-graphics-vide… · home › topics › ai-tools › article
[ARTICLE · art-146613] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

How to Render a Motion Graphics Video with Claude Opus 5.5: Prompt, Frame Renderer, ffmpeg and Real Cost

A developer documented a reproducible pipeline for turning Claude Opus 5.5 into a motion-graphics renderer by prompting the model to emit a self-contained HTML file with a deterministic renderFrame(t) function, then driving it frame-by-frame through a headless Chrome renderer and ffmpeg to produce a 12-second 1280x720 MP4. The approach requires the model to avoid Date.now, performance.now, Math.random and accumulated state so each frame is a pure function of time, and the writeup includes the exact prompt, a 25-line frame renderer, the ffmpeg command and the real API cost, run through the apimodels.app gateway.

by read6 min views1 publishedOct 7, 2026

Claude Opus 5.5 cannot output a single pixel, yet people keep posting motion-graphics videos "made by Opus 5.5". The trick is that the model writes a program that draws every frame, and you render that program into an MP4. This post is the full, reproducible pipeline we used for a 12-second 1280x720 promo: the exact prompt, the API call, a 25-line frame renderer, the ffmpeg command, and what it actually cost.

Disclosure up front: we ran this through apimodels.app, the API gateway I work on. apimodels.app is a multi-model API gateway: one API key and an OpenAI-compatible endpoint for about 150 image, video, audio and language models, including claude-opus-5-5. Every step below works the same against any OpenAI-compatible endpoint that serves Opus 5.5; only the base URL changes.

Ask Opus 5.5 for an HTML page with a deterministic renderFrame(t) function, not for "a video", then let a headless browser call that function once per frame.

Opus 5.5 outputs text. A page that animates itself with requestAnimationFrame looks fine in a browser, but you cannot export it frame-accurately, because what you see depends on timing. A page whose every frame is a pure function of t can be rendered frame 0, frame 1 … frame 359 in any order, and each call returns the same picture. That is exactly what a video encoder needs.

claude-opus-5-5 (OpenAI chat-completions format).puppeteer-core (npm i puppeteer-core works; we use the Chrome already installed on the machine). The rules in the first half of this prompt are what make the output renderable. The story comes second, the style last. This is the prompt we sent, unedited:

You are writing a program that renders a video. Output one complete, self-contained HTML file and nothing else.

Rules for the program (these make it renderable frame by frame):
1. A single <canvas> of exactly 1280x720, no CSS scaling, black page background.
2. Expose a global function window.renderFrame(t) that draws the frame at time t (seconds, 0 <= t < 12) from scratch. Everything on screen must be a pure function of t: no Date.now, no performance.now, no Math.random (use a seeded PRNG if you need noise), no accumulated state between calls, no requestAnimationFrame loop.
3. Also expose window.DURATION = 12 and window.FPS = 30. When the page is opened normally, play it once in a loop with requestAnimationFrame by calling renderFrame, so a human can preview it.
4. No external assets, fonts, images or libraries. Use system-ui for text.
5. The rhythm is 120 BPM: a beat every 0.5 s. Put every major cut, text entrance and impact on a beat time (a multiple of 0.5). List the beat plan as a comment at the top of the script.

The story (tell this, then decide the visuals):
A 12-second promo for "APIMODELS", an API gateway. Problem, then turn, then payoff:
- 0-3 s: a developer's screen is cluttered with many different API keys and SDK logos flying in from every side, overlapping, getting chaotic (draw them as simple rounded labels such as "image", "video", "LLM", "audio", "key_1", "key_2", "sdk", not real company logos).
- 3-4.5 s: everything snaps together on a beat and collapses into one glowing key.
- 4.5-9 s: from that one key, lines fan out to a grid of model cards that light up one per beat (label them "145 models", "image", "video", "LLM", "audio", "one endpoint").
- 9-12 s: clean end card: the word APIMODELS, the line "One key. Every model.", and "apimodels.app" small underneath; hold still for the last second.

Style: dark background, one accent colour (electric violet #7c5cff) plus white, motion with easing (easeOutCubic / easeInOutQuad written by you), subtle depth (scale and blur fall-off), no clutter in the final 3 seconds.

Before the code, do not explain. Return only the HTML.

Three details mattered more than the wording:

Date.now, performance.now, Math.random and accumulated state are the usual culprits. The model followed all four.

curl https://api.apimodels.app/v1/chat/completions \
  -H "Authorization: Bearer $APIMODELS_API_KEY" \
  -H "Content-Type: application/json" \
  -d "$(jq -n --rawfile p prompt.md \
        '{model:"claude-opus-5-5", stream:true, messages:[{role:"user", content:$p}]}')" \
  --no-buffer > stream.txt

Two things will bite you here if you skip them:

max_tokens. max_tokens: 4096 cap cuts the file off mid-script. If you omit the field, apimodels.app uses the model's own ceiling instead of inventing a small default. Join the delta.content pieces from stream.txt and save the result as promo.html. Open it in a browser first: it should play its own preview loop.

// render.mjs: pnpm add puppeteer-core (or npm i puppeteer-core)
import puppeteer from 'puppeteer-core'
import { mkdirSync, writeFileSync } from 'node:fs'

mkdirSync('frames', { recursive: true })
const browser = await puppeteer.launch({
  executablePath: '/Applications/Google Chrome.app/Contents/MacOS/Google Chrome',
})
const page = await browser.newPage()
await page.setViewport({ width: 1280, height: 720 })
await page.goto('file://' + process.cwd() + '/promo.html')
await page.evaluate(() => { window.requestAnimationFrame = () => 0 }) // stop the preview loop
const { DURATION, FPS } = await page.evaluate(() => ({ DURATION, FPS }))
for (let i = 0; i < DURATION * FPS; i++) {
  const png = await page.evaluate((t) => {
    window.renderFrame(t)
    return document.querySelector('canvas').toDataURL('image/png')
  }, i / FPS)
  writeFileSync(`frames/f${String(i).padStart(4, '0')}.png`, Buffer.from(png.split(',')[1], 'base64'))
}
await browser.close()

On Linux, point executablePath at your Chromium binary. 360 frames took 14 seconds on a laptop, and the page threw no errors.

ffmpeg -framerate 30 -i frames/f%04d.png -c:v libx264 -pix_fmt yuv420p -crf 20 -movflags +faststart promo.mp4

yuv420p and faststart are what make the file play everywhere, including in browsers and on phones. The result was a 1.8 MB, 12-second, 30 fps MP4.

Before judging the video, pull frames at the moments the brief describes and put them side by side: 2.3 s (the clutter), 3.5 s (one key), 7.2 s (cards lighting up) and 11.3 s (end card).

All four matched on the first attempt. To change something, ask by timecode and beat ("at 7.5 s the sixth card should light up with a pulse") and ask for the whole file back, so the frame contract stays intact.

Item Value
Model claude-opus-5-5
First token / total time 110 s / 6 min 37 s
Output tokens 31,786 (mostly thinking)
Cost at Anthropic list price ($4 / $20 per 1M tokens) about $0.64
Cost at apimodels.app list price ($2.40 / $12 per 1M tokens) about $0.38
Rendering and encoding local, free

Output tokens dominate. A clip that needs 30,000 to 80,000 tokens of thinking and code costs roughly $0.36 to $0.96 per attempt at the lower rate, so budget for two or three attempts per finished clip.

Code-rendered video is good at kinetic typography, UI walkthroughs, animated charts, logo reveals, explainers and social templates you can re-render with new text in seconds. It is the wrong tool for anything that has to look filmed: photoreal people, natural camera motion, skin, fabric, physics. For those shots use a text-to-video model and let Opus 5.5 write the shot list and prompts instead.

It also has no audio (add it with ffmpeg), and the model never runs its own page. Treat the HTML as untested code: render it and look at the frames before you trust it.

What would you render first with a renderFrame(t) contract: a chart, a logo reveal, or a UI walkthrough? I'm curious which kind of clip breaks the approach.

This post was written from our own run's code and numbers; AI helped with structure and editing.

── more in #ai-tools 4 stories · sorted by recency
── more on @claude opus 5.5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-render-a-moti…] indexed:0 read:6min 2026-10-07 · —