{"slug": "show-hn-demoforge-product-demo-videos-generated-from-playwright-scripts", "title": "Show HN: DemoForge – Product demo videos generated from Playwright scripts", "summary": "DemoForge, a new open-source tool, generates product demo videos from Playwright scripts by driving an app, recording it, auto-zooming on every click from the click log, drawing a synthetic cursor, speaking per-step narration and exporting an MP4. The tool requires Node, pnpm and ffmpeg, supports local espeak-ng narration or Gemini, ElevenLabs and Mistral Voxtral with API keys, and exports an 86-second demo in about 2 minutes via native ffmpeg or more slowly through ffmpeg.wasm in the browser. A companion MV3 Chrome extension records flows by hand via chrome.tabCapture, and a Claude Code skill writes demos/<name>/flow.json from a page.", "body_md": "Product demo videos from a script. Write the flow once as Playwright steps with a line of narration on each; DemoForge drives the app, records it, zooms in on every click, draws a smooth cursor, speaks the narration and exports an MP4. When the UI changes, run the same command again.\n\n- **Auto-zoom from the click log** , not computer vision — the browser already\nknows where you clicked.\n- **Synthetic cursor** with natural motion, since Playwright never moves a real one.\n- **Narration** : local espeak-ng, or Gemini / ElevenLabs / Mistral with a key.\nThe recorder holds each step until its line has been said.\n- **Editor** in the browser: drag zooms, cut, captions, background, then export.\n- A **Chrome extension** for recording by hand, feeding the same editor.\n\nNeeds Node, pnpm and ffmpeg. espeak-ng (`dnf install espeak-ng` /\n`apt install espeak-ng`) gives you a free local voice; no API key needed.\n\n```\npnpm install\npnpm -C packages/core build\npnpm exec playwright-core install chromium\n\nnode scripts/demoforge.mjs record demos/todomvc/flow.json   # drive the app, record\nnode scripts/demoforge.mjs open   demos/todomvc             # review in the editor\nnode scripts/demoforge.mjs export demos/todomvc             # -> demos/todomvc/demo.mp4\n```\n\n`demos/todomvc/flow.json` records Playwright's public TodoMVC demo — copy it\nand point `url` at your own app.\n\n```\npackages/core       @demoforge/core — types, zoom planner, evaluator, cursor path\napps/editor         Vite + React + Tailwind — player, timeline, compositor, export\napps/extension      MV3 Chrome extension — capture by hand + click log\nscripts/            demoforge.mjs — the record / open / export CLI\n```\n\n`pnpm -r test` runs the tests. `demoforge-docs/` has the product vision and\narchitecture notes.\n\nNarration needs a voice provider; nothing else does, and the editor tells you which ones are available.\n\n- \n**espeak-ng (local)** — free, offline, robotic.\n- \n**Gemini AI** — copy`.env.example` to`.env` and put a key in`GEMINI_API_KEY` . The same key writes the script. Sounds like a person.`.env` at the repo root or in`apps/editor/` both work; the key is read by the dev server only and never\nreaches the browser.A free-tier key allows only a few requests a minute, so generating a long script pauses when the quota says to and picks up again — the button tells you how long it is waiting.\n- \n**ElevenLabs** —`ELEVENLABS_API_KEY` , same two locations. The best voices.\nMetered per character: the free tier is 10,000 characters a month, personal\nuse only, and asks you to credit ElevenLabs. The panel shows what is left\nand what the next generate will cost, and warns before a run that would run\nout partway.\n- \n**Mistral Voxtral** —`MISTRAL_API_KEY` , same two locations.\n\n1. **Record.**`pnpm -C apps/extension build` , load`apps/extension/dist` unpacked in Chrome, open any`http(s)` page, click the DemoForge action →\nStart. It captures the tab with`chrome.tabCapture` and logs every click,\ninput, scroll and navigation against the recorder's own clock.\n2. **Stop.** Two files land in`~/Downloads/demoforge/<timestamp>/` :`recording.webm` and`demo.json` (a`DemoRecording` ).\n3. **Edit.**`pnpm -C apps/editor dev` and drop both into the editor. Zooms are\nplanned from the click log; drag the pills to move or resize them, add and\ndelete, restyle the frame.\n4. **Export.** Render to MP4. With the dev server running this uses native\nffmpeg (`apps/editor/vite-export.ts` ): the recording is decoded straight\nthrough, each frame drawn by the editor's own`compose()` on a Skia canvas,\nand piped into x264 — an 86 s demo in about 2 minutes. Without a server it\nfalls back to ffmpeg.wasm in the browser, several times slower.\n\nTell Claude Code \"record a demo of /login\" and it does the rest. The\n`demoforge` skill (`.claude/skills/demoforge/`, symlink it into\n`~/.claude/skills/` to use it from any repo) has the agent read the page,\nwrite `demos/<name>/flow.json` — Playwright actions with a `say` line of\nnarration on each — and run:\n\n```\nnode scripts/demoforge.mjs record demos/<name>/flow.json   # video + demo.json + project with script\nnode scripts/demoforge.mjs steps  demos/<name>             # frames the agent checks\nnode scripts/demoforge.mjs open   demos/<name>             # editor at ?demo=<name>\n```\n\nThe recorder paces itself to the narration: each line starts a beat before\nits action and the step holds until the line has been said. `setup` steps\n(signing in) run off camera. Capture is 2× device pixels, encoded as VP9 from the\nscreencast's own frames — Playwright's built-in recorder is capped at 1 Mbps. `${NAME}` in a flow is read from the environment\nor `.env`, so credentials stay out of the file. The flow is the only file in a\ntake that is committed — re-recording after a UI change is the same command.\n\nAn icon rail on the right opens six panels:\n\n| Panel | What it does | \n|---|---|\n| Script & voice | Narration lines on the timeline, drafted from the click log or written by hand, spoken by a local TTS | \n| Background | Image / Colour / Gradient tabs — 18 generated wallpapers, custom upload, gradient presets with editable stops and angle | \n| Zoom | Auto-zoom toggle, per-zoom or global scale, re-plan from the click log | \n| Captions | Add at the playhead, draft a set from the click log, edit text, position / size / colour | \n| Effects | Padding, corner radius, shadow blur / offset / strength | \n| Layout | Output aspect — Original, 16:9, 9:16, 1:1, 4:3, 4:5 | \n| Cursor | Show, click pulse, size, smoothing | \n\n**Aiming a zoom.** Select a pill and the preview drops back to the unzoomed\nframe with a rectangle showing exactly what that zoom will crop. Click or drag\nanywhere on the frame to move it, and a floating inspector gives you the zoom\nlevel, focus mode, reset and delete.\n\nA zoom's focus is **auto** by default: it points at the nearest click, and\nre-aims itself if you drag the pill somewhere else on the timeline. Placing a\npoint by hand switches it to **manual**, and nothing moves it again until you\nreset it.\n\nCaptions are drawn over the frame but outside the zoom transform — a caption\nbelongs to the viewer, not to the picture, so it does not slide or grow when\nthe camera moves. **From clicks** drafts one cue per click out of the element\ntext the extension already recorded; that is string formatting, not AI.\n\n**Cutting.** Press `T` to drop a cut at the playhead, then drag its edges;\n`I` and `O` cut everything before or after the playhead. Cut spans are shaded\nacross every lane, playback jumps over them, and they are gone from the export\n— video and audio both.\n\nCuts are the one place timeline time and source time come apart. The edit list\nin `packages/core/src/edits.ts` owns that conversion, and the rule that keeps\nit cheap is that **everything else stays in source time**: zoom keyframes,\ncaptions, the cursor path and the event log are never remapped. The exporter\nwalks edited time, maps each frame back through `editedToSource()`, and\ncomposites at a source timestamp exactly as the preview does.\n\nThe timeline has a scrubbable ruler with amber marks at every logged click, a\ncut lane, a zoom lane, a caption lane, and a clip lane. **Ctrl+Scroll** zooms the view\nabout the pointer, **Shift+Scroll** pans, and the window follows the playhead.\n\n| Key |  | \n|---|---|\n| `Space` | play / pause | \n| `Z` | add a zoom at the playhead | \n| `C` | add a caption at the playhead | \n| `N` | add a narration line at the playhead | \n| `S` | save the project | \n| `T` | cut a section out at the playhead | \n| `I`` O` | cut everything before / after the playhead | \n| `Delete` | remove the selection | \n| `Esc` | deselect | \n| `←``→` | step one frame (hold `Shift` for a second) | \n| `Home`` End` | jump to start / end | \n\nWallpapers are generated, not shipped — a base colour plus soft radial blobs, painted by one function used for both the picker swatch and the full frame, so the swatch cannot lie and there are no binary assets in the repo.\n\nA **script** is a list of lines, each anchored to a source timestamp on the\nsame clock as everything else. There are two ways to get one.\n\n**Write the script with AI** is the good one. The editor breaks the demo into\nsteps — an opening, then one per click — grabs a frame of the screen at each,\nrings the spot that was clicked, and sends the lot to a model along with how\nmany seconds it has to talk at each step. It writes to that budget, naming\nwhat is actually on screen. Give it a sentence about what the demo is for and\nit gets markedly better; that brief is saved with the project.\n\nTwo things the model is deliberately not trusted with:\n\n- **Timestamps.** It says which*step* a line belongs to; the editor decides\nwhen that lands. Asked for milliseconds, a model returns plausible ones, and\nplausible is not synchronised.\n- **Length.** Each step carries a word budget from its own window. An\nover-long line is trimmed back to a sentence boundary, never mid-sentence —\na line that runs a little long still reads, a truncated one does not.\n\nFrames of your recording go to Google when you press it. Nothing else in the editor sends anything anywhere.\n\n**From clicks** is the offline fallback: one line per step from the element\ntext already in the log, string templates, no network. Rough, but instant.\n\n**Generate voiceover** then speaks every line that has changed, measures how\nlong it actually took, and lays the results onto one track.\n\nThe mixdown plays in the preview (the captured tab audio stays muted there) and is muxed into the export with the recording ducked underneath the voice.\n\nA few things follow from how it is wired:\n\n- **Audio is saved beside the project, not inside it.****Save project** writes`<name>.dfp.json` and, when there is a voiceover,`<name>.narration.wav` . Drop both back in with the video and the narration\nplays immediately — nothing is respoken, which matters when the voice is\nmetered. The JSON stays a few readable kilobytes rather than megabytes of\nbase64, and the`.wav` is an ordinary file you can listen to or edit\nelsewhere.\n- **A reloaded mixdown is sliced back into lines.** Each line knows its anchor\nand how long it ran, so the track is cut up and put back in the speech\ncache. Change one line of a reloaded project and only that line is spoken\nagain. Drop the`.wav` and everything still works — it just costs a full\nregenerate.\n- **The mixdown is in source time** , so cuts splice it through the exact same\nfilter as the tab audio. Nothing in the narration path knows what a cut is.\n- **Level is applied once, at the end.** The mixdown is at unity; the preview\nsets it on the audio element and the export sets it with a filter, so the\ntwo agree and moving the slider never forces a re-mix.\n- **A line's length is its speech** , not something you drag. Until it has been\nspoken the timeline uses a word-count estimate, marked`est.` . If lines start\ntalking over each other the panel says so and offers to space them out.\n- **Providers live behind one endpoint.**`GET /api/tts` says who can speak\nand with which voices;`POST /api/tts` returns WAV\n(`apps/editor/vite-tts.ts` ). The panel renders whatever the server reports,\nso adding a provider is a server-side change. The mixdown, the timeline and\nthe exporter never learn who spoke.\n- **Different providers take different dials.** espeak-ng takes words per\nminute; Gemini takes a**director's note** — free text describing tone, pace\nand accent, handed to the model alongside the line; ElevenLabs takes neither\nand puts everything in the choice of voice. The panel shows only the dials\nthe chosen provider actually uses, and switching provider or note re-speaks\nthe affected lines.\n- **Voice lists come from the provider.** ElevenLabs' are fetched live against\nyour key, so your own cloned voices appear and no hardcoded id can go stale.\nIts availability means the key*works* , not just that one is set.\n- **Both hosted providers return raw PCM** at 24 kHz — Gemini describes it in\na mime type, ElevenLabs is asked for`pcm_24000` — so the WAV header is\nwritten server-side before the audio ever reaches the browser. (44.1 kHz\nfrom ElevenLabs needs a Pro subscription; 24 kHz does not.)\n- **\"Available\" means a key is configured** , not that it works — a bad key\nsurfaces as the provider's own error the first time you generate.\n\n**Save project** (or `S`) writes `<name>.dfp.json` — the recording plus every\nedit (zooms, captions, cuts, the narration script and the style) as plain\nreadable JSON — and `<name>.narration.wav` alongside it when there is a\nvoiceover. Drop the JSON, the video, and the `.wav` back in to carry on\nexactly where you were. A raw `demo.json` still opens too; it just gets\nfreshly planned zooms.\n\nThe schema lives in `packages/core/src/project.ts`, not in the editor, because\nthe point is that the editor is not the only thing that can write one. A\nscript, a CI job, or Claude Code can open a project, change the zooms or\ncaptions, write it back, and the editor will render exactly that.\n\n```\n{\n  \"format\": \"demoforge-project\",\n  \"version\": 1,\n  \"mediaName\": \"recording.webm\",       // referenced, not embedded\n  \"narrationName\": \"recording.narration.wav\",   // ditto; \"\" when there is none\n  \"brief\": \"AirSense is an air-quality dashboard for facilities teams.\",\n  \"recording\": { /* the DemoRecording from capture */ },\n  \"zooms\": [\n    { \"tStart\": 1500, \"tEnd\": 3700, \"targetXNorm\": 0.42, \"targetYNorm\": 0.31,\n      \"scale\": 1.8, \"easing\": \"easeInOutCubic\", \"focus\": \"auto\" }\n  ],\n  \"captions\": [\n    { \"tStart\": 2000, \"tEnd\": 4200, \"text\": \"Click \\\"Add Widget\\\"\" }\n  ],\n  \"script\": [                          // narration; text only, audio is regenerated\n    { \"tStart\": 800, \"text\": \"Start by clicking Add Widget.\", \"audioMs\": 2100 }\n  ],\n  \"style\": { \"background\": { \"kind\": \"wallpaper\", \"id\": \"cobalt\" }, \"aspect\": null,\n             \"padding\": 0.05, \"radius\": 0.02, \"shadow\": { \"blur\": 0.05, \"y\": 0.018, \"alpha\": 0.5 },\n             \"cursor\": { \"show\": true, \"size\": 0.045, \"smoothing\": 0.4, \"clicks\": true },\n             \"captions\": { \"size\": 0.045, \"position\": \"bottom\", \"color\": \"#ffffff\",\n                           \"background\": \"rgba(2,6,23,0.72)\" },\n             \"voice\": { \"provider\": \"local\", \"voice\": \"en-us+f3\", \"rate\": 170,\n                        \"direction\": \"\", \"gain\": 1, \"duck\": 0.25 } }\n}\n```\n\nNotes for anything editing one by hand:\n\n- `zooms` and`captions` are**time-ordered** , so \"the third zoom\" is stable.\nNeither list may overlap itself.\n- All coordinates are **0..1** , all style lengths are**fractions of the\noutput's shorter side** . No pixels anywhere.\n- `focus: \"auto\"` means the zoom is aimed at the nearest click and will re-aim\nif moved;`\"manual\"` pins it.\n- `parseProject()` is a trust boundary: it sorts and de-overlaps the lists,\nclamps every number into range, drops zero-length spans, and falls back to\ndefaults rather than letting`NaN` reach the renderer. It throws only on a\nmissing recording or a format version it does not understand — so a\nroughly-right file loads rather than failing.\n- `cuts` are spans of source video the demo skips, in source time. They are\nsorted, clamped, merged when they overlap, and dropped when shorter than\n100 ms. A legacy`trim: {startMs, endMs}` (a span to*keep* ) is migrated\ninto the equivalent head and tail cuts.\n- `script` lines carry`audioMs` only as a cached measurement. Change`text` and drop it — a stale length lays the timeline out for audio that no longer\nexists. Lines may overlap; that is reported, not prevented.\n- `style.voice.duck` is what the captured recording drops to while the voice\nis talking,`gain` is the voice's own level.\n- `brief` is what the demo is about, in your words. It is context for whoever\nwrites the narration — the AI writer reads it — and it is worth keeping so\nthe next rewrite starts from the same understanding.\n- `narrationName` names the rendered voiceover sitting next to the project.\nThe editor takes any dropped`.wav` as the narration, so the name is a hint\nrather than a requirement — files get renamed.\n- `style.voice.provider` is`\"local\"` ,`\"gemini\"` or`\"elevenlabs\"` ;`voice` is that`voice` is that provider's own id (`en-us+f3` ,`Iapetus` , or an ElevenLabs\nvoice id).`rate` is used by the local provider,`direction` by Gemini —\nboth are always stored, so switching provider and back keeps your settings.\n- Omitting `zooms` ,`captions` ,`cuts` ,`script` or`style` entirely is fine;\nthey default.\n\nEverything flows through one type, `DemoRecording`, defined once in\n`packages/core`. The rules that make Phase 1's editor reusable unchanged in\nPhase 2 (AI voiceover) and Phase 3 (Playwright re-record):\n\n- **Nothing downstream reads `source`.** Not the planner, editor, compositor or\nexporter. That branch is the seam that would break Phase 3.\n- **Stored coordinates are normalised 0..1.** Pixels appear at exactly one\nplace,`render/geometry.ts` , against the video's real decoded size.\n- **One clock.**`t0` is stamped at`MediaRecorder.start()` ; every`DemoEvent.t` is ms since it. The editor reconciles against the decoded\nduration on import and warns on drift over 100 ms.\n- **The editor draws its own cursor** from the click log and never depends on a\ncaptured OS cursor — which is what makes Phase 3 work for free.\n- **Audio is a separate track** , muxed at export, never an input to zoom logic.\n\n```\n./scripts/make-fixture.sh          # markers at exactly known coordinates\npnpm -C apps/editor dev            # drop fixture/demo.json + recording.webm\n```\n\nEvery zoom must land dead centre on its marker; the double-click at 5.0/5.12 s\nmust produce one zoom, and the giant panel at 13 s none. `pnpm -r test` asserts\nall of that headlessly.\n\nFor recording-side timing, see `apps/extension/README.md`.\n\n- **Velocity continuity at zoom seams.** Chained zooms are position-continuous,\nbut the ease curve's velocity still steps at a ramp boundary, which can read\nas a small jerk. A spring chasing the eased target is the fix if it shows.\n- **Native export is one process, one frame at a time** (~1.5x real time at\n2880x1800). Splitting the timeline across worker threads is the next step\nif that matters. The ffmpeg.wasm fallback still holds every frame in memory.\n- **Pill handles** stay inside the pill, so a very short zoom is fiddly to grab.\n- **Voice providers are dev-server only** , so a statically built editor has\nnone. Fine while this is a tool you run with`pnpm dev` ; the fix is the same\nbackend the Phase 2 TODO already calls for.\n- **Ducking is a constant** , not a sidechain compressor — right when the tab\naudio is ambience, wrong if it ever carries something worth hearing.", "url": "https://wpnews.pro/news/show-hn-demoforge-product-demo-videos-generated-from-playwright-scripts", "canonical_source": "https://github.com/DixitRam/Demo-Forge", "published_at": "2026-10-06 04:42:22+00:00", "updated_at": "2026-10-06 05:19:10.646349+00:00", "lang": "en", "topics": ["developer-tools", "ai-tools"], "entities": ["DemoForge", "Playwright", "espeak-ng", "Gemini", "ElevenLabs", "Mistral Voxtral", "Claude Code", "ffmpeg"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-demoforge-product-demo-videos-generated-from-playwright-scripts", "markdown": "https://wpnews.pro/news/show-hn-demoforge-product-demo-videos-generated-from-playwright-scripts.md", "text": "https://wpnews.pro/news/show-hn-demoforge-product-demo-videos-generated-from-playwright-scripts.txt", "jsonld": "https://wpnews.pro/news/show-hn-demoforge-product-demo-videos-generated-from-playwright-scripts.jsonld"}}