You type one prompt. "A man wakes up, walks to the window, and looks at the city. Anime style." You press enter.
And you get a video. A real one. For a moment it feels like magic.
Then you watch it again. The man in the bed has black hair. The man at the window has brown hair. The clock was round, and now it is gone. The floor was wood in the first second and carpet in the third.
It is an animation. It is also garbage. ๐๏ธ
Not because the model is bad. Because the model does not know your world โ your man, your clock, your room. Every second it draws them again from nothing, and every time a little different. You asked for one thing and got a hundred small guesses, glued together.
Your next idea is a stronger prompt. Describe the man, the clock, the room, every scene, every camera. Try it. The prompt grows to a page, then three, and the hair still changes. A prompt that really pins down one man, one room, three props and six scenes is not three pages. It is a book. No model reads a book and keeps all of it in mind for every frame.
So you stop asking for one thing. You cut the impossible job into eight small jobs, and at every step you hand the model something to copy instead of something to imagine.
An AI animation is many small calls, and after a week you cannot remember which prompt made which image. So every step is one markdown file of prompts next to one folder of results:
story/
โโโ story.md # 1 ๐ฌ the six lines, the worlds, characters, props
โโโ assets.md # 2 ๐จ one prompt per world, character, prop
โ # 3 ๐๏ธ โฆ and per raw scene
โโโ scenes.md # 4 ๐ one prompt per first/end frame
โ # 5 ๐๏ธ โฆ and per half-screen between scenes
โโโ video.md # 7 ๐ฅ one prompt per 4-second clip (6 ๐๏ธ the sound is in it)
โโโ film.txt # 8 โ๏ธ the edit: clip order and joins
โโโ cut.sh # 8 โ๏ธ builds the film from film.txt
โโโ assets/
โโโ scenes/
โโโ video/
โโโ final.mp4
Every prompt has the same three-line header, then the prompt:
| Header | Says |
|---|---|
| Output | the file this prompt makes |
| Attach | which earlier files go with it, and why |
| Use | where the result is needed later |
The markdown is the project; the images and clips are its output. Read a .md top to bottom and you see the film before it exists. Change one prompt, make one file again, and nothing else moves. Put the folder in git and you see what you changed last Tuesday.
Prompts in files, not in chat history. The chat is gone next week. The file is not.
Before any tool, you need a story. And it must be small โ not because small is beautiful, but because AI is bad at keeping things the same.
| Limit | Max | Why |
|---|---|---|
| ๐ Worlds | 2 | every place has its own light and colour |
| ๐ง Characters | 2 | each one must look the same in every scene |
| ๐งฉ Props | 3 | they repeat, so they must match |
| ๐๏ธ Scenes | 6 | every scene is a new chance to fail |
Everything else โ the wall, the sheet, the sky โ appears once and can change. Do not spend time on it.
Count what repeats. That is what the AI has to get right twice.
The story in this article: a man wakes up, walks to the window, and sees the city. Two worlds, one character, three props, six scenes:
| # | Scene |
|---|---|
| 1 | A man is asleep in his bed |
| 2 | The alarm clock on the bedside table rings |
| 3 | The man wakes up |
| 4 | He gets out of bed and walks to the window |
| 5 | He stands at the window and looks out |
| 6 | The city, as he sees it |
Write your six lines. List and count your worlds, characters and props. No prompts yet. No tools yet.
The next thing most people do is open an image model and type scene 1. Then scene 2, and the man is a different man.
Make assets first: one reference image of each world, character and prop, alone, on a plain background. Six assets, six prompts in assets.md. I use GPT Image 2.5 Sunburst, because it can look at attached images โ the whole method depends on that.
Four rules for every asset:
bed.png โ and they will not match.
The first prompt sets the style, and every later prompt copies this paragraph word for word. I will write [STYLE] for it from here on:
Flat cel colour fills, one hard shadow tone and a soft highlight,
crisp dark ink outlines, simplified shapes, limited palette,
hand-drawn anime TV series look. Not photorealistic, not a 3D render.
No text, no labels, no watermark.
๐ The world. Empty, with a fixed camera. You choose the camera once, and every scene in this room uses it.
## The room
- Output: assets/room.png
- Attach: nothing
- Use: background, scenes 1 to 5
Anime background illustration, 16:9, no characters, no furniture,
no objects. [STYLE]
A small empty bedroom in the early morning. Plain cream walls, a
wooden floor, a white ceiling. Pale morning light from the far wall
falls across the floor. No bed, no table, no window, no door.
Camera, fixed for this room: standing eye height, from the door,
looking toward the far wall. Left wall and far wall both visible,
the floor filling the bottom third of the frame.
๐ง The character. A sheet, not a picture: at least three views, full body from head to feet, nothing in the hands. In scene 4 the man walks across the room, and the model has to know his feet. If the sheet stops at the chest, the model guesses, and it guesses differently every time.
## The man
- Output: assets/man.png
- Attach: assets/room.png (colours and light only, do not draw the room)
- Use: character reference, every scene
Character reference sheet, 16:9, plain light-grey background.
The exact art style of the attached image. [STYLE]
Character only: no props, nothing in the hands, no background.
Three views side by side, same scale, each the full body from head
to feet, nothing cropped: front, side, three-quarter. Same face,
clothes and colours in all three.
A man around thirty, slim, light skin. Short messy black hair, tired
brown eyes, a day of stubble. Plain white T-shirt, grey pyjama
trousers, bare feet. Front: arms at his sides, eyes half open.
Side: standing straight. Three-quarter: one hand rubbing the back
of his neck, a small yawn.
๐งฉ The props. The opposite: one view, in full detail. The viewer knows the man by his face. The viewer knows the clock by its bells, its red body, its black numbers. So name every part.
## The clock
- Output: assets/clock.png
- Attach: assets/room.png (colours and light only)
- Use: prop reference, scenes 1 to 3
Prop reference sheet, 16:9, plain light-grey background, no
characters, no hands. The exact art style of the attached image.
[STYLE]
One view only, large in the frame: three-quarter from slightly
above, so the face and the top are both visible.
A round red alarm clock, old style. Red metal body with a soft shine.
Two silver bells on top, a small silver hammer between them, a silver
ring handle behind. White face, black numbers 1 to 12, black hour and
minute hands, a thin red second hand. Two short black legs. No glow,
no digital display.
Do the same for the bed and the window. Then the city, with nothing attached and its own fixed camera: from the window, looking out.
One prompt, one thing, alone. Scenes come later, and they only copy.
This is the cheapest place to be wrong: a bad asset costs one image. Fix the man's hair here and it is right in all six scenes. Open all the assets side by side โ same film, same size โ and do not move on until they match.
Each of the six lines becomes one image. They go in assets/ too, numbered โ raw material for the frames in Step 4.
assets/
โโโ room.png โฆ window.png
โโโ 01-asleep.png
โโโ 02-alarm.png
โโโ 03-awake.png
โโโ 04-walk.png
โโโ 05-window.png
โโโ 06-city.png
Spend a minute on the names, because the same name travels through every step: 02-alarm.png โ 02-alarm-first.png โ 02-alarm.mp4. Two digits first, so files sort in story order. One word after, the thing the viewer sees. Lowercase, no spaces. When you are twenty files deep and a clip looks wrong, 04-walk tells you which prompt to open. IMG_0417 tells you nothing.
The prompts live in scenes.md. A scene attaches everything that appears in it:
## 02 ยท The alarm
- Output: assets/02-alarm.png
- Attach: assets/room.png ยท assets/man.png ยท assets/bed.png ยท assets/clock.png
Single illustration, 16:9, in the exact 2D anime style of the
attached images. [STYLE]
The attached images are the only source of truth. room.png is the
room: same walls, floor, light and camera. man.png is the man: same
face, hair and clothes. bed.png is the bed and clock.png is the clock,
exactly as drawn. Draw nothing that is not in the attached images or
described below. Only the poses and the action change.
Camera: the fixed camera of the room, from the door.
The bed against the left wall, the man asleep in it, on his side,
eyes closed. The bedside table next to it, the clock on it, ringing:
bells blurred with motion, three small motion lines on each side.
Look at how little of this is about the scene. Style, copied. The camera, copied. A list of what each image is. The action is four lines.
Something I took a while to accept: image models read pictures better than words. Write "a round red alarm clock with two silver bells" and you get a different clock every time. Attach clock.png and say "this clock", and you get that clock. When a scene is hard, do not reach for a longer prompt. Reach for another picture.
| Rule | Why |
|---|---|
| โ Attach only what appears | the city is not in scene 2, so city.png stays out โ extra images confuse it |
| ๐ข Attach in order | world, then characters, then props โ the first image is the base |
| ๐ท๏ธ Name every attachment | "room.png is the room" โ the model does not know which picture is which |
| ๐ท Same camera as the world | five scenes from one camera look like a film; from five cameras, a mess |
A scene prompt describes the action. The pictures describe everything else.
๐ The comic-strip test. When all six are done, put them in a row and read them like a comic strip with no words. If the pictures tell the story by themselves, your story and your scenes are right. If you reach one and think "wait, what happened here?", a scene is missing or shows the wrong moment. Do not fix it in the next step. Go back to the six lines, change them, and make that scene again. A hole here becomes a hole in the film.
Here is the tricky part. Look at scene 2. The clock is ringing. Now imagine the clip. Is this image the first frame or the last?
It is the last. The clip starts with a quiet clock, then it rings.
A video model that works from images wants two: where the clip starts and where it ends. So every scene needs two frames, and you already have one. Decide which, then copy it into a new scenes/ folder with the answer as a suffix. The raw scene stays in assets/.
| Scene | What you have | It is | Still needed |
|---|---|---|---|
| 01 ยท asleep | the man asleep, the room still | first | end: he turns over in his sleep |
| 02 ยท alarm | the clock ringing | end | first: the clock still |
| 03 ยท awake | the man sitting up, eyes open | end | first: eyes closed, head on the pillow |
| 04 ยท walk | the man standing by the bed | first | end: the man at the window, his back to us |
| 05 ยท window | the man at the window | first | end: the same, the curtain moved by the wind |
| 06 ยท city | the city, wide | first | end: the same city, the camera a little closer |
scenes/
โโโ 01-asleep-first.png
โโโ 02-alarm-end.png
โโโ 03-awake-end.png
โโโ 04-walk-first.png
โโโ 05-window-first.png
โโโ 06-city-first.png
Half the files are missing. To make each one, attach the frame you already have โ the finished scene itself โ plus only the assets involved in the change. The prompt is tiny, because you describe one difference:
## 04 ยท The walk, end frame
- Output: scenes/04-walk-end.png
- Attach: scenes/04-walk-first.png ยท assets/man.png ยท assets/window.png
Single illustration, 16:9, the same style as the attached scene.
04-walk-first.png is the frame this picture follows: same room,
camera, bed and light. man.png is the man, for his face, hair and
clothes. window.png is the window, exactly as drawn.
One change only: the man has crossed the room. He stands at the
window on the far wall, his back to the camera, one hand on the
curtain. The bed is empty, the blanket pushed back. Everything else
stays exactly where it is.
The man moved, so his sheet is attached again, so his back is right. He touches the window now, so it is attached. The room and the bed come from the scene itself.
You do not describe a scene twice. You describe it once, then describe what changed.
Put the clips in a row and watch. Scene 1 ends with the man turning in his sleep. Scene 2 starts with him still. Same room, but the arm moved, the blanket moved, and your eye catches the jump. Six scenes, five jumps.
My first fix was a clip for the gap itself: from the end of scene 1 to the first frame of scene 2, so nothing would ever cut. I spent a lot of time on this. It does not work. The model has to invent motion between two frames that were never meant to connect, and what it invents is a slow, strange morph. It looks worse than the jump.
Two honest choices:
A half-screen is one frame, no first and no end. Attach the world for the light, the character or prop it shows, and write a small prompt. Name it after the scene it follows, with -half:
## 02 ยท half-screen, the eyes
- Output: scenes/02-alarm-half.png
- Attach: assets/room.png (light only) ยท assets/man.png
Single illustration, 16:9, the exact style of the attached images.
[STYLE] man.png is the man: same face, hair and stubble.
Extreme close-up of the man's face, filling the frame, on his side on
the pillow, eyes closed. Morning light across his face from the
right. One eyebrow slightly raised, as if the ringing has just
reached him. No bed edge, no clock, no room.
| After scene | Half-screen | Small motion in the clip |
|---|---|---|
| 01 ยท asleep | the clock face, close | the second hand ticks |
| 02 ยท alarm | the man's closed eyes | the eyebrow lifts |
| 03 ยท awake | bare feet touching the wooden floor | the toes curl |
| 04 ยท walk | his hand on the white curtain | the curtain sways |
| 05 ยท window | his eyes, open, with light in them | a slow blink |
You do not need all five. Use one where the jump is ugly, a plain cut where it is not.
In the edit, a half-screen dissolves in and out, half a second to a second on each side, over the scene before and the scene after. So a 4-second half-screen shows alone for two to three seconds. The eye is on the close-up while the room changes underneath it โ that is what makes the jump disappear.
A half-screen is a cut that looks like it was planned.
You may want to make the sound now โ a voice from a voice model, the alarm, the city โ and give it to the video model with the frames. You cannot. Not with the two frames.
Seedance 2.5 runs on many platforms. I use it inside ElevenLabs โ the same place I would make the voice โ and even there, a voice file and a first-and-end frame pair cannot go into the same request. On fal.ai there is no audio input at all. On ByteDance's own API you can attach an audio reference, but the moment you do, the first and end frames stop being first and end โ they become loose references, and the clip no longer runs from one to the other. I tried this more than once. Each time I got a good audio file and no place for it.
So the voice goes in the prompt. Write the line in quotes, describe the voice, say who speaks and when. The model renders the voice over the clip, with the mouth on the words, and makes the room sound too โ the ring, the sheets, the far city. I have tested this many times. It is not perfect, but it is good, and it sits exactly where the picture needs it, because the same model made both.
He stands at the window and says, in a low, tired voice, a man in
his thirties just awake: "Morning." His mouth moves with the word.
Only he speaks.
Sound: his voice, the curtain, the city far below. No music.
Honest about quality: ElevenLabs' own voice model is better. But I cannot attach it next to my two frames, so it does not matter how good it is. Maybe a future version will take both. Until then, the best voice is the one you can actually put in the clip.
The voice you cannot attach is not a voice. It is a file.
One rule from here: ๐ต no music in the clips. Music goes over the whole film at the end, in one piece.
One clip per scene and per half-screen, into video/, prompts in video.md. For a scene, attach the first and the end frame. For a half-screen, the single frame.
Three rules, and they all say keep it short:
| Keep short | How | Why |
|---|---|---|
| โฑ๏ธ The clip | 4 seconds | the model is at its best in short clips; long ones drift |
| ๐ The motion | one thing moves | the man walks, or the clock rings โ not both |
| โ๏ธ The prompt | a few lines | a long prompt makes worse motion, not better |
Why 4 seconds and not less? Because Seedance will not go lower. Many moments are shorter than that โ a clock starts ringing in one second โ but the clip is 4 seconds whether you need them or not. So put the motion in the middle of the clip and leave the first and last second quiet: still at the start, hold at the end. Those quiet seconds are what the edit fades over in Step 8. If the action starts on frame one, the fade eats it.
## 04 ยท The walk
- Output: video/04-walk.mp4
- Attach: scenes/04-walk-first.png (first) ยท scenes/04-walk-end.png (end) ยท 4 s
Image-to-video, 4 seconds, from the first frame to the end frame.
Keep the room, camera, man and bed exactly as drawn.
0-1 s: he stands by the bed, still. 1-3 s: he walks slowly to the
window, bare feet on wood. 3-4 s: he stops, back to us, his hand
reaches for the curtain. Hold the end frame.
Sound: soft footsteps on wood. No music.
## 04 ยท half-screen, the curtain
- Output: video/04-walk-half.mp4
- Attach: scenes/04-walk-half.png (first frame only) ยท 4 s
Image-to-video, 4 seconds, from this single frame; the picture holds
to the end. Small motion only: the curtain sways once, the fingers
tighten on the cloth. No camera move.
Sound: the curtain, the city far away. No music.
The temptation is to add the light, the mood, what the man feels. Every line you add, the model obeys by making the motion worse. Say what moves, say when, say what it sounds like, and stop.
Short clip, one motion, few words. The frames do the talking.
You will remake some clips. When one is wrong, check the frames first โ if the two frames do not agree, no prompt saves the clip. If the frames are right, cut the prompt, do not grow it.
I use ffmpeg. It is free, it runs anywhere, and the edit becomes a text file you can read next month.
๐ The order. One clip per line, and after each name, how that clip comes in: cut or fade. A scene after a scene is a cut. Anything touching a half-screen is a fade. This file is the edit. Save it as film.txt:
01-asleep cut
01-asleep-half fade
02-alarm fade
02-alarm-half fade
03-awake fade
03-awake-half fade
04-walk fade
04-walk-half fade
05-window fade
05-window-half fade
06-city fade
๐ The join. A fade overlaps two clips by FADE seconds, picture and sound together โ one second is soft, half a second keeps more of the half-screen on screen. A cut is a one-frame crossfade: invisible, but it removes the click a hard cut leaves in the sound. This script reads the list and builds the chain:
#!/bin/sh
set -e
LEN=4; FADE=1
set --
for n in $(awk '{print $1}' film.txt); do set -- "$@" -i "video/$n.mp4"; done
N=$(wc -l < film.txt | tr -d ' ')
FILTER=$(awk -v len="$LEN" -v fade="$FADE" '
{ n++; join[n] = $2 }
END {
v = "[0:v]"; a = "[0:a]"; t = len
for (i = 2; i <= n; i++) {
d = (join[i] == "fade") ? fade : 0.04
printf "%s[%d:v]xfade=transition=fade:duration=%s:offset=%.2f[v%d];", v, i-1, d, t - d, i
printf "%s[%d:a]acrossfade=d=%s[a%d];", a, i-1, d, i
v = "[v" i "]"; a = "[a" i "]"; t += len - d
}
}' film.txt)
ffmpeg -y "$@" -filter_complex "${FILTER%;}" -map "[v$N]" -map "[a$N]" \
-c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac film.mp4
Run sh cut.sh. Eleven clips of 4 seconds, ten fades of one second: 34 seconds of film, no clicks. Change a line in film.txt โ swap a fade for a cut, drop a half-screen โ and run it again.
๐ต The finish. A fade from black and to black, then the music over the whole film, quietly under the clips' own sound:
ffmpeg -y -i film.mp4 \
-vf "fade=t=in:d=0.5,fade=t=out:st=33.5:d=0.5" \
-af "afade=t=in:d=0.5,afade=t=out:st=33.5:d=0.5" \
film-faded.mp4
ffmpeg -y -i film-faded.mp4 -i music.mp3 \
-filter_complex "[1:a]volume=0.2,afade=t=out:st=30:d=4[m];[0:a][m]amix=inputs=2:duration=first" \
-c:v copy final.mp4
The edit is a text file. Change one line, run it again.
Two models, and you can count every call before you start. List prices, October 2026: Seedance 2.5 at 720p with sound is $0.473 per second on fal.ai (other platforms differ). OpenAI bills GPT Image 2.5 by tokens and publishes no per-image price; the closest published figure is $0.165 for a 1536ร1024 high-quality image. Run one call, read the usage field, and put your own number in.
| Step | Model | What | Calls | Each | Cost |
|---|---|---|---|---|---|
| 2 ๐จ | GPT Image 2.5 | 2 worlds, 1 character, 3 props | 6 | $0.17 | $0.99 |
| 3 ๐๏ธ | GPT Image 2.5 | 6 raw scenes | 6 | $0.17 | $0.99 |
| 4 ๐ | GPT Image 2.5 | 6 second frames | 6 | $0.17 | $0.99 |
| 5 ๐๏ธ | GPT Image 2.5 | 5 half-screens | 5 | $0.17 | $0.83 |
| 7 ๐ฅ | Seedance 2.5 | 11 clips ร 4 s, at $0.473 / s | 11 | $1.89 | $20.81 |
| 34 seconds of film | 34 | $24.61 |
About twenty-five dollars, if every call comes out right the first time. It will not. Plan for every second call needing a retry, and the film costs closer to forty.
All twenty-three images together cost about as much as two clips. So the expensive mistake is never a bad image โ it is a bad frame you only notice once it moves. That is why every step before 7 ends with "do not move on until they match." Spend the cheap calls. Save the expensive ones.
One folder, five text files, and a film at the bottom. When something is wrong โ and something will be โ you know which file to open.
first or end, and its partner is made from it
If you have been further than this โ a model that takes a voice reference, a better way to bridge two scenes โ tell me in the comments. I am still looking for the fix to the one step I could not make work.