A new account went from zero to 18,971 views. Here is how each film was made, what it cost, and where the views came from.
18,971 views in
5 days across Instagram, YouTube Shorts and TikTok, since the first post on 27 September
Platform-counted views, not people. Read from each platform about every half hour and drawn every three hours. Hover or drag across the chart for any moment.
Last week we started a short-video channel with no followers and no ad budget. Two coding agents, Claude Code and Codex, ran the production: script, narration, images, animation, edit, checks and upload. We made twelve films about famous frauds, styled like handmade clay, and posted two or three a day.
Every image and video clip on fal.ai, including every take we threw away, cost $184. Narration, music and the first three films’ video came from an ElevenLabs startup grant. In five days the films drew 18,971 views and 25 followers. For a week’s work and $184, that was more than I expected.
This is the whole thing: one film from start to finish, the ledger, what seemed to work, and the one thing I can't explain.
The twelve #
Nine true stories of fraud and imposture, and three office comedies about Sam, a tired analyst. Hover to play one with sound, or tap on a phone. Views are totals across all platforms. Every post is labelled as AI-made; real people appear as clay caricatures in dramatized stories.
One film, start to finish #
Every film is a folder with one film.json in it. The agents generate everything else from that file. Here is the Holmes film, stage by stage.
- The line### Start with the first sentenceThe script is data: beats of narration, and under each beat the shots, with a prompt and a motion for each. The first line has to be one concrete, strange act with its object on screen. Here: one drop of blood. { "id": "s01", "vo": "She raised seven hundred million dollars on one drop of blood.", "shots": [{ "id": "h01", "refs": ["holmes"], "prompt": "Close-up: Elizabeth Holmes in her black turtleneck holds up between two fingers a tiny glossy clear clay vial with one bright red drop of blood…", "motion": "she lifts the tiny vial toward camera and the red drop glints; she does not blink" }] }
- The voice### Narrate a beat at a timeEach beat is a separate text-to-speech call, then transcribed back with Whisper and compared to the script. Whole-script narration once quietly changed a line, so now nothing goes forward until the words match. “She raised seven hundred million dollars on one drop of blood.” 3.3 seconds · ElevenLabs v4 · each word checked against Whisper before anything else is made
- The cast### Draw each character onceOne reference sheet per character. Every shot she appears in is generated with this sheet attached, so her face, hair and turtleneck stay the same across the film.
- The frame### One still per shotAn image model (Nano Banana, $0.039 an image) draws each shot from its prompt and the cast sheet. Every prompt starts with the same paragraph about plasticine, fingerprints, felt, warm light and a tilt-shift look. That paragraph does most of the art direction.
Photorealistic macro photograph of a handmade stop-motion plasticine claymation miniature. All characters sculpted from matte plasticine with visible fingerprints and sculpting-tool marks, big glossy white bead eyes with heavy sculpted clay eyelids, clothes made of felt and knitted wool, miniature handmade set, warm cinematic lighting, shallow depth of field, tilt-shift miniature look. Everything is clay or felt, no real human skin, no real hands. No text, no letters, no writing anywhere. Vertical 9:16 composition.
- The motion### Five seconds of movementEach still becomes a five-second clip. This film used Kling 2.5 Turbo Pro; after testing ten video models on the same frames we moved every shot to Kling 3 Pro, at $0.112 a second. Prompts ask for small, physical motion: she lifts the vial, the drop glints, she doesn't blink.
- The cut### Cut to the voiceShots are cut on frame counts against the narration’s word timings, captions are drawn two to four words at a time, and the score is composed to the film’s exact length. Press Sound on to hear it.
- The check### Look at every shotEight frames from every shot of the exact file we upload, plus four frames a second of every clip. A script then checks pace, loudness, the length of the ending and more. Posting refuses a film that fails.
- The post### Post it everywhereOne file goes to Instagram, TikTok, YouTube Shorts and Facebook at a fixed time, labelled as AI-made. Then we read the numbers before writing the next one. Holmes so far 1,607 Instagram160 TikTok3 YouTube Shorts0 Facebook
Making it look handmade #
The look is one paragraph prepended to every image prompt: matte plasticine with visible fingerprints and tool marks, glossy bead eyes under heavy lids, clothes of felt and knitted wool, warm light, tilt-shift. No text anywhere, because image models still can't spell. Before writing it we studied every video on @clayrified, a clay channel we admire.
Photorealistic macro photograph of a handmade stop-motion plasticine claymation miniature. All characters sculpted from matte plasticine with visible fingerprints and sculpting-tool marks, big glossy white bead eyes with heavy sculpted clay eyelids, clothes made of felt and knitted wool, miniature handmade set, warm cinematic lighting, shallow depth of field, tilt-shift miniature look. Everything is clay or felt, no real human skin, no real hands. No text, no letters, no writing anywhere. Vertical 9:16 composition.
Three rules did more than any prompt trick. Every character gets a reference sheet attached to every shot. Shots with nobody in them say the scene is empty of any living thing
; no people
on its own produced invented creatures. And every acting element has to be in the still, because the video model draws in whatever the still leaves out.
The video model mattered most. We tested ten of them. Here are nine on the same shot, at the same moment. Kling 3 Pro gave the best character acting: watch the hands and the eyes. One test clip per model is a thin basis for a choice, so we treated it as provisional and checked it on whole films.
What it cost #
Generation was the cheap part. A finished film now costs $7 to $11 to generate. The surprise was the agents themselves. We ran them on flat subscriptions, but at API prices Claude Code, run as one long conversation, would have spent up to $47 of tokens per film, mostly re-reading its own history. Priced that way, the coding agent costs more than the video model.
| For the twelve films | Paid |
|---|---|
| Images and video on fal.ai, including rejected takes (list price of every logged request, all week) | $184 |
| Voice, music, and the first three films’ images and video on ElevenLabs (startup grant; $112 at list price) | $0 |
| Claude Code on a Claude Max 20x plan ($200 a month) and Codex on a ChatGPT Pro plan, subscriptions we already pay for all our engineering | no extra |
| Cash spent on the films | $184 |
About one cent of cash per view. Had we paid API prices instead of the subscriptions, the agents’ tokens would have added up to about $268 and the ElevenLabs work $112, for about three cents a view. The fal.ai and ElevenLabs figures cover the whole week, including drafts, a model comparison and films not shown here, so they overstate the twelve slightly. Not counted: my own time reviewing cuts.
Some of it was waste we caused. A fifth of the generation budget went on takes we rejected. On one evening the agents made three explainer films twice over before I had seen a single frame, and $61 of finished work went in the bin. The fix is boring: a person approves one frame and one clip before any batch. Here is what each film cost at list prices:
SWEAT AI FILM SUPPLY
Elizabeth Holmes
TOTAL**$7.14**
Every film, itemised from our own logs, including the takes we threw away.
The first three films ran their video through ElevenLabs before we moved to fal.ai, which is why they cost several times more. The coding agents run on monthly subscriptions we pay for anyway, so a film adds nothing to that bill.
Where the views came from #
Almost two-thirds came from Instagram. TikTok gave every film between 30 and 250 views, whatever we did. Facebook gave 11 views in total.
12,054 Instagram
5,005 YouTube Shorts
1,901 TikTok
11 Facebook
The first line seems to matter more than anything else. Below are the twelve opening lines, ranked by the share of Instagram viewers still watching after three seconds. The top five each name one strange, physical thing in the first sentence: a press printing money, a drop of blood, a fake prince’s dinner. The bottom of the list is the three comedies and the films that opened on a wide shot or a setup. That is twelve films with different stories and posting times, not a test, so read it as a pattern rather than a proof.
And one thing I can't explain. Five of our first seven films got 740 to 1,398 views on YouTube Shorts. Every film since Holmes, on the afternoon of 30 September, has had between zero and four. There are no notices or strikes in YouTube Studio, the videos are public, and nothing changed in how we upload. Instagram didn't drop on the same days. If you have seen YouTube stop testing a new channel's Shorts like this, I would like to hear what it was.
What the agents were good at #
They were good at everything that can be measured. Each time I found a problem, the agent wrote it into the pipeline as a check, usually within the hour: sample more frames, cap the length of the ending, refuse narration that was sped up afterwards. Each check kept that particular mistake from shipping again.
Every film teaches the next one
Each film changes one thing. After it posts, we read each platform about every half hour, find the second where people stopped watching, and turn that into a rule the next film has to pass before it can post.
- 1Post One change per film
- 2Measure Each platform, about every half hour
- 3Find the leak The second where viewers leave
- 4Write the rule The next film must pass it to post
Six things the numbers taught us in three days. The rule cards show They sent Sam a pizza menu for proof of address, film 11: the shaded band is what the rule allows, the dot is where that film landed.
- 28 Sep after film 2#### The opening shotEiffel opened on a wide shot, with the hook sentence split over two cuts. Mavrodi opened on a close-up of the act, the hook sentence over one 3-second shot. TikTok: still watching at 3 secondsMavrodi, close-up56% Eiffel, wide shot41% Instagram: average watchMavrodi, close-up37 s Eiffel, wide shot18 s Instagram skip rate: 23% for Mavrodi, 45% for Eiffel. Once the scorecard was live, Eiffel held 49% and Mavrodi 76%. What changedEvery film opens on one concrete act in close-up, and the hook stays on screen long enough to read. The rule nowHook on screenpass2.9 s Rule: at least 2.5 sLong enough to read the opening object. Wide establishing shots lost viewers in the first second.
- 28 Sep same review#### The pitch at the endThe 20-second Sweat pitch at the end lost 60% of the people still watching. About 2% reached the last frame. What changedThe product ending is capped at 6 seconds, from the first mention of Sweat to the last frame. An earlier film ran 11.4. The rule nowProduct endingpass5.9 s was 11.4Rule: at most 6 sFrom the first mention of Sweat to the last frame. Long pitches lost most of the people still watching.
- 28 Sep same evening#### Endings that looked brokenThe music was composed to a guessed length and chopped where the film ended, and the end card was squeezed to about 3 seconds. Both read as glitches. What changedMusic is composed to the film’s real length, and the silence after the last word is capped. One film had 4.5 seconds of it. The rule nowSilence after the last wordpass0.45 s was 4.5Rule: at most 0.75 sDead air at the end reads as a broken video.
- 29 Sep after five films#### The first lineThe winners’ first lines named one concrete, absurd act: “printed his own money”, “double your money. With postage stamps.” Madoff’s first line was abstract, “clients almost never lost money”, and did as badly as Eiffel’s wide shot. TikTok: still watching at about 4 secondsBig Bull58% Ponzi54% Mavrodi53% Eiffel36% Madoff33% All twelve first lines are ranked above .What changedThe first sentence names one thing someone did, not a summary of who they were. The rule nowFirst lineruleOne concrete act in the first sentence.
- 29 Sep TikTok Studio#### The first frame on TikTokIn the feed, the first frame was our logo overlay, a still object and a year as the first caption. It read as the start of an ad. TikTok Studio said most viewers stopped watching at 0:01. What changedThe brand moved later in the film, and the first caption became the hook phrase instead of the year. The rule nowFirst frameruleThe brand comes later. The first caption is the hook phrase, not a year.
- 30 Sep still open#### The second lineWith the opening fixed, TikTok kept 82 to 92% of viewers at one second, so the leak moved to seconds 3 to 10. A second line that adds a new object or a turn kept far more of them than one that gives a name and background. TikTok: of those watching at 3 s, still there at 10 s (about)New object or turn70% Name and background40% New object or turn: Gignac, Mehta, Mavrodi. Name and background: Holmes, Madoff. What changedNot a rule yet. Five films is a pattern, not a proof, so new films are now assigned one kind of second line at random. The rule nowSecond linetestingRandomized test, decided at 13 films per arm.
Five rules came from making the films, not from the numbers
pass
3.22 words/s
Natural speed from the voice model itself. Speeding it up afterwards made audible artefacts.
pass
2.2 s
Cut on the narration’s clause breaks, like the clay channel we studied.
pass
23.2 dB
The score ducks under every word, but never disappears.
pass
−13.1 LUFS
Loud enough for a phone speaker without clipping.
pass
96 frames
Eight from every shot of the exact file we upload.
The pipeline reads views at 24, 48 and 72 hours, the share still watching after three seconds, average watch time, shares and follows. Each film is judged on each platform against the median of the films before it.
They were no good at taste. Every rejection of a style, a cast or a boring story came from a person watching, never from the agents' own review. Claude Code built the pipeline and made all twelve films. Codex later re-cut two of them to shorten their endings, reusing the footage it already had, which was the cheapest improvement of the week.
It is the same split we use in our day job: automated checks catch what can be measured, and a person signs off on the judgement. That is how our analysts review a business for KYB.