AI Animation From Idea to Film: Eight Small Jobs Instead of One Impossible Prompt A developer outlines an eight-step pipeline for AI animation that replaces a single monolithic prompt with small, file-based jobs, each handing the model reference images to copy rather than scenes to imagine. The approach caps a story at two worlds, two characters, three props and six scenes, and stores every prompt in markdown files alongside its output folder so individual steps can be regenerated without disturbing the rest. The author uses GPT Image 2.5 Sunburst to generate the reference assets. You type one prompt. "A man wakes up, walks to the window, and looks at the city. Anime style." You press enter. And you get a video. A real one. For a moment it feels like magic. Then you watch it again. The man in the bed has black hair. The man at the window has brown hair. The clock was round, and now it is gone. The floor was wood in the first second and carpet in the third. It is an animation. It is also garbage. ๐Ÿ—‘๏ธ Not because the model is bad. Because the model does not know your world โ€” your man, your clock, your room. Every second it draws them again from nothing, and every time a little different. You asked for one thing and got a hundred small guesses, glued together. Your next idea is a stronger prompt. Describe the man, the clock, the room, every scene, every camera. Try it. The prompt grows to a page, then three, and the hair still changes. A prompt that really pins down one man, one room, three props and six scenes is not three pages. It is a book. No model reads a book and keeps all of it in mind for every frame. So you stop asking for one thing. You cut the impossible job into eight small jobs, and at every step you hand the model something to copy instead of something to imagine. An AI animation is many small calls, and after a week you cannot remember which prompt made which image. So every step is one markdown file of prompts next to one folder of results : story/ โ”œโ”€โ”€ story.md 1 ๐ŸŽฌ the six lines, the worlds, characters, props โ”œโ”€โ”€ assets.md 2 ๐ŸŽจ one prompt per world, character, prop โ”‚ 3 ๐ŸŽž๏ธ โ€ฆ and per raw scene โ”œโ”€โ”€ scenes.md 4 ๐Ÿ” one prompt per first/end frame โ”‚ 5 ๐Ÿ‘๏ธ โ€ฆ and per half-screen between scenes โ”œโ”€โ”€ video.md 7 ๐ŸŽฅ one prompt per 4-second clip 6 ๐ŸŽ™๏ธ the sound is in it โ”œโ”€โ”€ film.txt 8 โœ‚๏ธ the edit: clip order and joins โ”œโ”€โ”€ cut.sh 8 โœ‚๏ธ builds the film from film.txt โ”œโ”€โ”€ assets/ โ”œโ”€โ”€ scenes/ โ”œโ”€โ”€ video/ โ””โ”€โ”€ final.mp4 Every prompt has the same three-line header, then the prompt: | Header | Says | |---|---| | Output | the file this prompt makes | | Attach | which earlier files go with it, and why | | Use | where the result is needed later | The markdown is the project; the images and clips are its output. Read a .md top to bottom and you see the film before it exists. Change one prompt, make one file again, and nothing else moves. Put the folder in git and you see what you changed last Tuesday. Prompts in files, not in chat history. The chat is gone next week. The file is not. Before any tool, you need a story. And it must be small โ€” not because small is beautiful, but because AI is bad at keeping things the same. | Limit | Max | Why | |---|---|---| | ๐ŸŒ Worlds | 2 | every place has its own light and colour | | ๐Ÿง Characters | 2 | each one must look the same in every scene | | ๐Ÿงฉ Props | 3 | they repeat, so they must match | | ๐ŸŽž๏ธ Scenes | 6 | every scene is a new chance to fail | Everything else โ€” the wall, the sheet, the sky โ€” appears once and can change. Do not spend time on it. Count what repeats. That is what the AI has to get right twice. The story in this article: a man wakes up, walks to the window, and sees the city. Two worlds, one character, three props, six scenes: | | Scene | |---|---| | 1 | A man is asleep in his bed | | 2 | The alarm clock on the bedside table rings | | 3 | The man wakes up | | 4 | He gets out of bed and walks to the window | | 5 | He stands at the window and looks out | | 6 | The city, as he sees it | Write your six lines. List and count your worlds, characters and props. No prompts yet. No tools yet. The next thing most people do is open an image model and type scene 1. Then scene 2, and the man is a different man. Make assets first: one reference image of each world, character and prop, alone, on a plain background. Six assets, six prompts in assets.md . I use GPT Image 2.5 Sunburst, because it can look at attached images โ€” the whole method depends on that. Four rules for every asset: bed.png โ€” and they will not match. The first prompt sets the style, and every later prompt copies this paragraph word for word. I will write STYLE for it from here on: Flat cel colour fills, one hard shadow tone and a soft highlight, crisp dark ink outlines, simplified shapes, limited palette, hand-drawn anime TV series look. Not photorealistic, not a 3D render. No text, no labels, no watermark. ๐ŸŒ The world. Empty, with a fixed camera. You choose the camera once, and every scene in this room uses it. The room - Output: assets/room.png - Attach: nothing - Use: background, scenes 1 to 5 Anime background illustration, 16:9, no characters, no furniture, no objects. STYLE A small empty bedroom in the early morning. Plain cream walls, a wooden floor, a white ceiling. Pale morning light from the far wall falls across the floor. No bed, no table, no window, no door. Camera, fixed for this room: standing eye height, from the door, looking toward the far wall. Left wall and far wall both visible, the floor filling the bottom third of the frame. ๐Ÿง The character. A sheet, not a picture: at least three views, full body from head to feet , nothing in the hands. In scene 4 the man walks across the room, and the model has to know his feet. If the sheet stops at the chest, the model guesses, and it guesses differently every time. The man - Output: assets/man.png - Attach: assets/room.png colours and light only, do not draw the room - Use: character reference, every scene Character reference sheet, 16:9, plain light-grey background. The exact art style of the attached image. STYLE Character only: no props, nothing in the hands, no background. Three views side by side, same scale, each the full body from head to feet, nothing cropped: front, side, three-quarter. Same face, clothes and colours in all three. A man around thirty, slim, light skin. Short messy black hair, tired brown eyes, a day of stubble. Plain white T-shirt, grey pyjama trousers, bare feet. Front: arms at his sides, eyes half open. Side: standing straight. Three-quarter: one hand rubbing the back of his neck, a small yawn. ๐Ÿงฉ The props. The opposite: one view, in full detail. The viewer knows the man by his face. The viewer knows the clock by its bells, its red body, its black numbers. So name every part. The clock - Output: assets/clock.png - Attach: assets/room.png colours and light only - Use: prop reference, scenes 1 to 3 Prop reference sheet, 16:9, plain light-grey background, no characters, no hands. The exact art style of the attached image. STYLE One view only, large in the frame: three-quarter from slightly above, so the face and the top are both visible. A round red alarm clock, old style. Red metal body with a soft shine. Two silver bells on top, a small silver hammer between them, a silver ring handle behind. White face, black numbers 1 to 12, black hour and minute hands, a thin red second hand. Two short black legs. No glow, no digital display. Do the same for the bed and the window. Then the city, with nothing attached and its own fixed camera: from the window, looking out. One prompt, one thing, alone. Scenes come later, and they only copy. This is the cheapest place to be wrong: a bad asset costs one image. Fix the man's hair here and it is right in all six scenes. Open all the assets side by side โ€” same film, same size โ€” and do not move on until they match. Each of the six lines becomes one image. They go in assets/ too, numbered โ€” raw material for the frames in Step 4. assets/ โ”œโ”€โ”€ room.png โ€ฆ window.png โ”œโ”€โ”€ 01-asleep.png โ”œโ”€โ”€ 02-alarm.png โ”œโ”€โ”€ 03-awake.png โ”œโ”€โ”€ 04-walk.png โ”œโ”€โ”€ 05-window.png โ””โ”€โ”€ 06-city.png Spend a minute on the names, because the same name travels through every step: 02-alarm.png โ†’ 02-alarm-first.png โ†’ 02-alarm.mp4 . Two digits first, so files sort in story order. One word after, the thing the viewer sees. Lowercase, no spaces. When you are twenty files deep and a clip looks wrong, 04-walk tells you which prompt to open. IMG 0417 tells you nothing. The prompts live in scenes.md . A scene attaches everything that appears in it: 02 ยท The alarm - Output: assets/02-alarm.png - Attach: assets/room.png ยท assets/man.png ยท assets/bed.png ยท assets/clock.png Single illustration, 16:9, in the exact 2D anime style of the attached images. STYLE The attached images are the only source of truth. room.png is the room: same walls, floor, light and camera. man.png is the man: same face, hair and clothes. bed.png is the bed and clock.png is the clock, exactly as drawn. Draw nothing that is not in the attached images or described below. Only the poses and the action change. Camera: the fixed camera of the room, from the door. The bed against the left wall, the man asleep in it, on his side, eyes closed. The bedside table next to it, the clock on it, ringing: bells blurred with motion, three small motion lines on each side. Look at how little of this is about the scene. Style, copied. The camera, copied. A list of what each image is. The action is four lines. Something I took a while to accept: image models read pictures better than words. Write "a round red alarm clock with two silver bells" and you get a different clock every time. Attach clock.png and say "this clock", and you get that clock. When a scene is hard, do not reach for a longer prompt. Reach for another picture. | Rule | Why | |---|---| | โœ… Attach only what appears | the city is not in scene 2, so city.png stays out โ€” extra images confuse it | | ๐Ÿ”ข Attach in order | world, then characters, then props โ€” the first image is the base | | ๐Ÿท๏ธ Name every attachment | "room.png is the room" โ€” the model does not know which picture is which | | ๐Ÿ“ท Same camera as the world | five scenes from one camera look like a film; from five cameras, a mess | A scene prompt describes the action. The pictures describe everything else. ๐Ÿ“– The comic-strip test. When all six are done, put them in a row and read them like a comic strip with no words. If the pictures tell the story by themselves, your story and your scenes are right. If you reach one and think "wait, what happened here?", a scene is missing or shows the wrong moment. Do not fix it in the next step. Go back to the six lines, change them, and make that scene again. A hole here becomes a hole in the film. Here is the tricky part. Look at scene 2. The clock is ringing. Now imagine the clip. Is this image the first frame or the last ? It is the last. The clip starts with a quiet clock, then it rings. A video model that works from images wants two: where the clip starts and where it ends. So every scene needs two frames, and you already have one. Decide which, then copy it into a new scenes/ folder with the answer as a suffix. The raw scene stays in assets/ . | Scene | What you have | It is | Still needed | |---|---|---|---| | 01 ยท asleep | the man asleep, the room still | first | end: he turns over in his sleep | | 02 ยท alarm | the clock ringing | end | first: the clock still | | 03 ยท awake | the man sitting up, eyes open | end | first: eyes closed, head on the pillow | | 04 ยท walk | the man standing by the bed | first | end: the man at the window, his back to us | | 05 ยท window | the man at the window | first | end: the same, the curtain moved by the wind | | 06 ยท city | the city, wide | first | end: the same city, the camera a little closer | scenes/ โ”œโ”€โ”€ 01-asleep-first.png โ”œโ”€โ”€ 02-alarm-end.png โ”œโ”€โ”€ 03-awake-end.png โ”œโ”€โ”€ 04-walk-first.png โ”œโ”€โ”€ 05-window-first.png โ””โ”€โ”€ 06-city-first.png Half the files are missing. To make each one, attach the frame you already have โ€” the finished scene itself โ€” plus only the assets involved in the change. The prompt is tiny, because you describe one difference: 04 ยท The walk, end frame - Output: scenes/04-walk-end.png - Attach: scenes/04-walk-first.png ยท assets/man.png ยท assets/window.png Single illustration, 16:9, the same style as the attached scene. 04-walk-first.png is the frame this picture follows: same room, camera, bed and light. man.png is the man, for his face, hair and clothes. window.png is the window, exactly as drawn. One change only: the man has crossed the room. He stands at the window on the far wall, his back to the camera, one hand on the curtain. The bed is empty, the blanket pushed back. Everything else stays exactly where it is. The man moved, so his sheet is attached again, so his back is right. He touches the window now, so it is attached. The room and the bed come from the scene itself. You do not describe a scene twice. You describe it once, then describe what changed. Put the clips in a row and watch. Scene 1 ends with the man turning in his sleep. Scene 2 starts with him still. Same room, but the arm moved, the blanket moved, and your eye catches the jump. Six scenes, five jumps. My first fix was a clip for the gap itself: from the end of scene 1 to the first frame of scene 2, so nothing would ever cut. I spent a lot of time on this. It does not work. The model has to invent motion between two frames that were never meant to connect, and what it invents is a slow, strange morph. It looks worse than the jump. Two honest choices: A half-screen is one frame, no first and no end. Attach the world for the light, the character or prop it shows, and write a small prompt. Name it after the scene it follows, with -half : 02 ยท half-screen, the eyes - Output: scenes/02-alarm-half.png - Attach: assets/room.png light only ยท assets/man.png Single illustration, 16:9, the exact style of the attached images. STYLE man.png is the man: same face, hair and stubble. Extreme close-up of the man's face, filling the frame, on his side on the pillow, eyes closed. Morning light across his face from the right. One eyebrow slightly raised, as if the ringing has just reached him. No bed edge, no clock, no room. | After scene | Half-screen | Small motion in the clip | |---|---|---| | 01 ยท asleep | the clock face, close | the second hand ticks | | 02 ยท alarm | the man's closed eyes | the eyebrow lifts | | 03 ยท awake | bare feet touching the wooden floor | the toes curl | | 04 ยท walk | his hand on the white curtain | the curtain sways | | 05 ยท window | his eyes, open, with light in them | a slow blink | You do not need all five. Use one where the jump is ugly, a plain cut where it is not. In the edit, a half-screen dissolves in and out, half a second to a second on each side, over the scene before and the scene after. So a 4-second half-screen shows alone for two to three seconds. The eye is on the close-up while the room changes underneath it โ€” that is what makes the jump disappear. A half-screen is a cut that looks like it was planned. You may want to make the sound now โ€” a voice from a voice model, the alarm, the city โ€” and give it to the video model with the frames. You cannot. Not with the two frames. Seedance 2.5 runs on many platforms. I use it inside ElevenLabs โ€” the same place I would make the voice โ€” and even there, a voice file and a first-and-end frame pair cannot go into the same request. On fal.ai there is no audio input at all. On ByteDance's own API you can attach an audio reference, but the moment you do, the first and end frames stop being first and end โ€” they become loose references, and the clip no longer runs from one to the other. I tried this more than once. Each time I got a good audio file and no place for it. So the voice goes in the prompt. Write the line in quotes, describe the voice, say who speaks and when. The model renders the voice over the clip, with the mouth on the words, and makes the room sound too โ€” the ring, the sheets, the far city. I have tested this many times. It is not perfect, but it is good, and it sits exactly where the picture needs it, because the same model made both. He stands at the window and says, in a low, tired voice, a man in his thirties just awake: "Morning." His mouth moves with the word. Only he speaks. Sound: his voice, the curtain, the city far below. No music. Honest about quality: ElevenLabs' own voice model is better. But I cannot attach it next to my two frames, so it does not matter how good it is. Maybe a future version will take both. Until then, the best voice is the one you can actually put in the clip. The voice you cannot attach is not a voice. It is a file. One rule from here: ๐ŸŽต no music in the clips. Music goes over the whole film at the end, in one piece. One clip per scene and per half-screen, into video/ , prompts in video.md . For a scene, attach the first and the end frame. For a half-screen, the single frame. Three rules, and they all say keep it short : | Keep short | How | Why | |---|---|---| | โฑ๏ธ The clip | 4 seconds | the model is at its best in short clips; long ones drift | | ๐Ÿƒ The motion | one thing moves | the man walks, or the clock rings โ€” not both | | โœ๏ธ The prompt | a few lines | a long prompt makes worse motion, not better | Why 4 seconds and not less? Because Seedance will not go lower. Many moments are shorter than that โ€” a clock starts ringing in one second โ€” but the clip is 4 seconds whether you need them or not. So put the motion in the middle of the clip and leave the first and last second quiet: still at the start, hold at the end. Those quiet seconds are what the edit fades over in Step 8. If the action starts on frame one, the fade eats it. 04 ยท The walk - Output: video/04-walk.mp4 - Attach: scenes/04-walk-first.png first ยท scenes/04-walk-end.png end ยท 4 s Image-to-video, 4 seconds, from the first frame to the end frame. Keep the room, camera, man and bed exactly as drawn. 0-1 s: he stands by the bed, still. 1-3 s: he walks slowly to the window, bare feet on wood. 3-4 s: he stops, back to us, his hand reaches for the curtain. Hold the end frame. Sound: soft footsteps on wood. No music. 04 ยท half-screen, the curtain - Output: video/04-walk-half.mp4 - Attach: scenes/04-walk-half.png first frame only ยท 4 s Image-to-video, 4 seconds, from this single frame; the picture holds to the end. Small motion only: the curtain sways once, the fingers tighten on the cloth. No camera move. Sound: the curtain, the city far away. No music. The temptation is to add the light, the mood, what the man feels. Every line you add, the model obeys by making the motion worse. Say what moves, say when, say what it sounds like, and stop. Short clip, one motion, few words. The frames do the talking. You will remake some clips. When one is wrong, check the frames first โ€” if the two frames do not agree, no prompt saves the clip. If the frames are right, cut the prompt, do not grow it. I use ffmpeg . It is free, it runs anywhere, and the edit becomes a text file you can read next month. ๐Ÿ“‹ The order. One clip per line, and after each name, how that clip comes in: cut or fade . A scene after a scene is a cut. Anything touching a half-screen is a fade. This file is the edit. Save it as film.txt : 01-asleep cut 01-asleep-half fade 02-alarm fade 02-alarm-half fade 03-awake fade 03-awake-half fade 04-walk fade 04-walk-half fade 05-window fade 05-window-half fade 06-city fade ๐Ÿ”— The join. A fade overlaps two clips by FADE seconds, picture and sound together โ€” one second is soft, half a second keeps more of the half-screen on screen. A cut is a one-frame crossfade: invisible, but it removes the click a hard cut leaves in the sound. This script reads the list and builds the chain: bash /bin/sh cut.sh โ€” joins the clips in film.txt into film.mp4 Every clip is LEN seconds. A fade overlaps two clips by FADE seconds. set -e LEN=4; FADE=1 set -- for n in $ awk '{print $1}' film.txt ; do set -- "$@" -i "video/$n.mp4"; done N=$ wc -l < film.txt | tr -d ' ' FILTER=$ awk -v len="$LEN" -v fade="$FADE" ' { n++; join n = $2 } END { v = " 0:v "; a = " 0:a "; t = len for i = 2; i <= n; i++ { d = join i == "fade" ? fade : 0.04 printf "%s %d:v xfade=transition=fade:duration=%s:offset=%.2f v%d ;", v, i-1, d, t - d, i printf "%s %d:a acrossfade=d=%s a%d ;", a, i-1, d, i v = " v" i " "; a = " a" i " "; t += len - d } }' film.txt ffmpeg -y "$@" -filter complex "${FILTER%;}" -map " v$N " -map " a$N " \ -c:v libx264 -crf 18 -pix fmt yuv420p -c:a aac film.mp4 Run sh cut.sh . Eleven clips of 4 seconds, ten fades of one second: 34 seconds of film, no clicks. Change a line in film.txt โ€” swap a fade for a cut , drop a half-screen โ€” and run it again. ๐ŸŽต The finish. A fade from black and to black, then the music over the whole film, quietly under the clips' own sound: ffmpeg -y -i film.mp4 \ -vf "fade=t=in:d=0.5,fade=t=out:st=33.5:d=0.5" \ -af "afade=t=in:d=0.5,afade=t=out:st=33.5:d=0.5" \ film-faded.mp4 ffmpeg -y -i film-faded.mp4 -i music.mp3 \ -filter complex " 1:a volume=0.2,afade=t=out:st=30:d=4 m ; 0:a m amix=inputs=2:duration=first" \ -c:v copy final.mp4 The edit is a text file. Change one line, run it again. Two models, and you can count every call before you start. List prices, October 2026: Seedance 2.5 at 720p with sound is $0.473 per second on fal.ai other platforms differ . OpenAI bills GPT Image 2.5 by tokens and publishes no per-image price; the closest published figure is $0.165 for a 1536ร—1024 high-quality image. Run one call, read the usage field, and put your own number in. | Step | Model | What | Calls | Each | Cost | |---|---|---|---|---|---| | 2 ๐ŸŽจ | GPT Image 2.5 | 2 worlds, 1 character, 3 props | 6 | $0.17 | $0.99 | | 3 ๐ŸŽž๏ธ | GPT Image 2.5 | 6 raw scenes | 6 | $0.17 | $0.99 | | 4 ๐Ÿ” | GPT Image 2.5 | 6 second frames | 6 | $0.17 | $0.99 | | 5 ๐Ÿ‘๏ธ | GPT Image 2.5 | 5 half-screens | 5 | $0.17 | $0.83 | | 7 ๐ŸŽฅ | Seedance 2.5 | 11 clips ร— 4 s, at $0.473 / s | 11 | $1.89 | $20.81 | | | | 34 seconds of film | 34 | | $24.61 | About twenty-five dollars, if every call comes out right the first time. It will not. Plan for every second call needing a retry, and the film costs closer to forty. All twenty-three images together cost about as much as two clips. So the expensive mistake is never a bad image โ€” it is a bad frame you only notice once it moves. That is why every step before 7 ends with "do not move on until they match." Spend the cheap calls. Save the expensive ones. One folder, five text files, and a film at the bottom. When something is wrong โ€” and something will be โ€” you know which file to open. first or end , and its partner is made from it If you have been further than this โ€” a model that takes a voice reference, a better way to bridge two scenes โ€” tell me in the comments. I am still looking for the fix to the one step I could not make work.