{"slug": "ai-animation-from-idea-to-film-eight-small-jobs-instead-of-one-impossible-prompt", "title": "AI Animation From Idea to Film: Eight Small Jobs Instead of One Impossible Prompt", "summary": "A developer outlines an eight-step pipeline for AI animation that replaces a single monolithic prompt with small, file-based jobs, each handing the model reference images to copy rather than scenes to imagine. The approach caps a story at two worlds, two characters, three props and six scenes, and stores every prompt in markdown files alongside its output folder so individual steps can be regenerated without disturbing the rest. The author uses GPT Image 2.5 Sunburst to generate the reference assets.", "body_md": "You type one prompt. \"A man wakes up, walks to the window, and looks at the city. Anime style.\" You press enter.\n\nAnd you get a video. A real one. For a moment it feels like magic.\n\nThen you watch it again. The man in the bed has black hair. The man at the window has brown hair. The clock was round, and now it is gone. The floor was wood in the first second and carpet in the third.\n\nIt is an animation. It is also garbage. 🗑️\n\nNot because the model is bad. Because the model does not know your world — your man, your clock, your room. Every second it draws them again from nothing, and every time a little different. You asked for one thing and got a hundred small guesses, glued together.\n\nYour next idea is a stronger prompt. Describe the man, the clock, the room, every scene, every camera. Try it. The prompt grows to a page, then three, and the hair still changes. A prompt that really pins down one man, one room, three props and six scenes is not three pages. It is a book. No model reads a book and keeps all of it in mind for every frame.\n\nSo you stop asking for one thing. You cut the impossible job into eight small jobs, and at every step you hand the model something to **copy** instead of something to imagine.\n\nAn AI animation is many small calls, and after a week you cannot remember which prompt made which image. So every step is **one markdown file of prompts next to one folder of results**:\n\n```\nstory/\n├── story.md      # 1 🎬 the six lines, the worlds, characters, props\n├── assets.md     # 2 🎨 one prompt per world, character, prop\n│                 # 3 🎞️ … and per raw scene\n├── scenes.md     # 4 🔁 one prompt per first/end frame\n│                 # 5 👁️ … and per half-screen between scenes\n├── video.md      # 7 🎥 one prompt per 4-second clip (6 🎙️ the sound is in it)\n├── film.txt      # 8 ✂️ the edit: clip order and joins\n├── cut.sh        # 8 ✂️ builds the film from film.txt\n├── assets/\n├── scenes/\n├── video/\n└── final.mp4\n```\n\nEvery prompt has the same three-line header, then the prompt:\n\n| Header | Says | \n|---|---|\n| **Output** | the file this prompt makes | \n| **Attach** | which earlier files go with it, and why | \n| **Use** | where the result is needed later | \n\nThe markdown is the project; the images and clips are its output. Read a `.md` top to bottom and you see the film before it exists. Change one prompt, make one file again, and nothing else moves. Put the folder in git and you see what you changed last Tuesday.\n\nPrompts in files, not in chat history. The chat is gone next week. The file is not.\n\nBefore any tool, you need a story. And it must be small — not because small is beautiful, but because AI is bad at keeping things the same.\n\n| Limit | Max | Why | \n|---|---|---|\n| 🌍 **Worlds** | 2 | every place has its own light and colour | \n| 🧍 **Characters** | 2 | each one must look the same in every scene | \n| 🧩 **Props** | 3 | they repeat, so they must match | \n| 🎞️ **Scenes** | 6 | every scene is a new chance to fail | \n\nEverything else — the wall, the sheet, the sky — appears once and can change. Do not spend time on it.\n\nCount what repeats. That is what the AI has to get right twice.\n\nThe story in this article: a man wakes up, walks to the window, and sees the city. Two worlds, one character, three props, six scenes:\n\n| # | Scene | \n|---|---|\n| 1 | A man is asleep in his bed | \n| 2 | The alarm clock on the bedside table rings | \n| 3 | The man wakes up | \n| 4 | He gets out of bed and walks to the window | \n| 5 | He stands at the window and looks out | \n| 6 | The city, as he sees it | \n\nWrite your six lines. List and count your worlds, characters and props. No prompts yet. No tools yet.\n\nThe next thing most people do is open an image model and type scene 1. Then scene 2, and the man is a different man.\n\nMake **assets** first: one reference image of each world, character and prop, alone, on a plain background. Six assets, six prompts in `assets.md`. I use GPT Image 2.5 Sunburst, because it can look at attached images — the whole method depends on that.\n\nFour rules for every asset:\n\n`bed.png` — and they will not match.\nThe first prompt sets the style, and every later prompt copies this paragraph word for word. I will write `[STYLE]` for it from here on:\n\n```\nFlat cel colour fills, one hard shadow tone and a soft highlight,\ncrisp dark ink outlines, simplified shapes, limited palette,\nhand-drawn anime TV series look. Not photorealistic, not a 3D render.\nNo text, no labels, no watermark.\n```\n\n🌍 **The world.** Empty, with a fixed camera. You choose the camera once, and every scene in this room uses it.\n\n```\n## The room\n- Output: assets/room.png\n- Attach: nothing\n- Use: background, scenes 1 to 5\n\nAnime background illustration, 16:9, no characters, no furniture,\nno objects. [STYLE]\n\nA small empty bedroom in the early morning. Plain cream walls, a\nwooden floor, a white ceiling. Pale morning light from the far wall\nfalls across the floor. No bed, no table, no window, no door.\n\nCamera, fixed for this room: standing eye height, from the door,\nlooking toward the far wall. Left wall and far wall both visible,\nthe floor filling the bottom third of the frame.\n```\n\n🧍 **The character.** A sheet, not a picture: at least three views, **full body from head to feet**, nothing in the hands. In scene 4 the man walks across the room, and the model has to know his feet. If the sheet stops at the chest, the model guesses, and it guesses differently every time.\n\n```\n## The man\n- Output: assets/man.png\n- Attach: assets/room.png (colours and light only, do not draw the room)\n- Use: character reference, every scene\n\nCharacter reference sheet, 16:9, plain light-grey background.\nThe exact art style of the attached image. [STYLE]\n\nCharacter only: no props, nothing in the hands, no background.\nThree views side by side, same scale, each the full body from head\nto feet, nothing cropped: front, side, three-quarter. Same face,\nclothes and colours in all three.\n\nA man around thirty, slim, light skin. Short messy black hair, tired\nbrown eyes, a day of stubble. Plain white T-shirt, grey pyjama\ntrousers, bare feet. Front: arms at his sides, eyes half open.\nSide: standing straight. Three-quarter: one hand rubbing the back\nof his neck, a small yawn.\n```\n\n🧩 **The props.** The opposite: one view, in full detail. The viewer knows the man by his face. The viewer knows the clock by its bells, its red body, its black numbers. So name every part.\n\n```\n## The clock\n- Output: assets/clock.png\n- Attach: assets/room.png (colours and light only)\n- Use: prop reference, scenes 1 to 3\n\nProp reference sheet, 16:9, plain light-grey background, no\ncharacters, no hands. The exact art style of the attached image.\n[STYLE]\n\nOne view only, large in the frame: three-quarter from slightly\nabove, so the face and the top are both visible.\n\nA round red alarm clock, old style. Red metal body with a soft shine.\nTwo silver bells on top, a small silver hammer between them, a silver\nring handle behind. White face, black numbers 1 to 12, black hour and\nminute hands, a thin red second hand. Two short black legs. No glow,\nno digital display.\n```\n\nDo the same for the bed and the window. Then the city, with nothing attached and its own fixed camera: from the window, looking out.\n\nOne prompt, one thing, alone. Scenes come later, and they only copy.\n\nThis is the cheapest place to be wrong: a bad asset costs one image. Fix the man's hair here and it is right in all six scenes. Open all the assets side by side — same film, same size — and do not move on until they match.\n\nEach of the six lines becomes one image. They go in `assets/` too, numbered — raw material for the frames in Step 4.\n\n```\nassets/\n├── room.png … window.png\n├── 01-asleep.png\n├── 02-alarm.png\n├── 03-awake.png\n├── 04-walk.png\n├── 05-window.png\n└── 06-city.png\n```\n\nSpend a minute on the names, because the same name travels through every step: `02-alarm.png` → `02-alarm-first.png` → `02-alarm.mp4`. Two digits first, so files sort in story order. One word after, the thing the viewer sees. Lowercase, no spaces. When you are twenty files deep and a clip looks wrong, `04-walk` tells you which prompt to open. `IMG_0417` tells you nothing.\n\nThe prompts live in `scenes.md`. A scene attaches everything that appears in it:\n\n```\n## 02 · The alarm\n- Output: assets/02-alarm.png\n- Attach: assets/room.png · assets/man.png · assets/bed.png · assets/clock.png\n\nSingle illustration, 16:9, in the exact 2D anime style of the\nattached images. [STYLE]\n\nThe attached images are the only source of truth. room.png is the\nroom: same walls, floor, light and camera. man.png is the man: same\nface, hair and clothes. bed.png is the bed and clock.png is the clock,\nexactly as drawn. Draw nothing that is not in the attached images or\ndescribed below. Only the poses and the action change.\n\nCamera: the fixed camera of the room, from the door.\n\nThe bed against the left wall, the man asleep in it, on his side,\neyes closed. The bedside table next to it, the clock on it, ringing:\nbells blurred with motion, three small motion lines on each side.\n```\n\nLook at how little of this is about the scene. Style, copied. The camera, copied. A list of what each image is. The action is four lines.\n\nSomething I took a while to accept: image models read pictures better than words. Write \"a round red alarm clock with two silver bells\" and you get a different clock every time. Attach `clock.png` and say \"this clock\", and you get that clock. When a scene is hard, do not reach for a longer prompt. Reach for another picture.\n\n| Rule | Why | \n|---|---|\n| ✅ **Attach only what appears** | the city is not in scene 2, so `city.png` stays out — extra images confuse it | \n| 🔢 **Attach in order** | world, then characters, then props — the first image is the base | \n| 🏷️ **Name every attachment** | \"room.png is the room\" — the model does not know which picture is which | \n| 📷 **Same camera as the world** | five scenes from one camera look like a film; from five cameras, a mess | \n\nA scene prompt describes the action. The pictures describe everything else.\n\n📖 **The comic-strip test.** When all six are done, put them in a row and read them like a comic strip with no words. If the pictures tell the story by themselves, your story and your scenes are right. If you reach one and think \"wait, what happened here?\", a scene is missing or shows the wrong moment. Do not fix it in the next step. Go back to the six lines, change them, and make that scene again. A hole here becomes a hole in the film.\n\nHere is the tricky part. Look at scene 2. The clock is ringing. Now imagine the clip. Is this image the **first** frame or the **last**?\n\nIt is the last. The clip starts with a quiet clock, then it rings.\n\nA video model that works from images wants two: where the clip starts and where it ends. So every scene needs two frames, and you already have one. Decide which, then copy it into a new `scenes/` folder with the answer as a suffix. The raw scene stays in `assets/`.\n\n| Scene | What you have | It is | Still needed | \n|---|---|---|---|\n| 01 · asleep | the man asleep, the room still | **first** | end: he turns over in his sleep | \n| 02 · alarm | the clock ringing | **end** | first: the clock still | \n| 03 · awake | the man sitting up, eyes open | **end** | first: eyes closed, head on the pillow | \n| 04 · walk | the man standing by the bed | **first** | end: the man at the window, his back to us | \n| 05 · window | the man at the window | **first** | end: the same, the curtain moved by the wind | \n| 06 · city | the city, wide | **first** | end: the same city, the camera a little closer | \n\n```\nscenes/\n├── 01-asleep-first.png\n├── 02-alarm-end.png\n├── 03-awake-end.png\n├── 04-walk-first.png\n├── 05-window-first.png\n└── 06-city-first.png\n```\n\nHalf the files are missing. To make each one, attach the frame you already have — the finished scene itself — plus only the assets involved in the change. The prompt is tiny, because you describe one difference:\n\n```\n## 04 · The walk, end frame\n- Output: scenes/04-walk-end.png\n- Attach: scenes/04-walk-first.png · assets/man.png · assets/window.png\n\nSingle illustration, 16:9, the same style as the attached scene.\n04-walk-first.png is the frame this picture follows: same room,\ncamera, bed and light. man.png is the man, for his face, hair and\nclothes. window.png is the window, exactly as drawn.\n\nOne change only: the man has crossed the room. He stands at the\nwindow on the far wall, his back to the camera, one hand on the\ncurtain. The bed is empty, the blanket pushed back. Everything else\nstays exactly where it is.\n```\n\nThe man moved, so his sheet is attached again, so his back is right. He touches the window now, so it is attached. The room and the bed come from the scene itself.\n\nYou do not describe a scene twice. You describe it once, then describe what changed.\n\nPut the clips in a row and watch. Scene 1 ends with the man turning in his sleep. Scene 2 starts with him still. Same room, but the arm moved, the blanket moved, and your eye catches the jump. Six scenes, five jumps.\n\nMy first fix was a clip for the gap itself: from the end of scene 1 to the first frame of scene 2, so nothing would ever cut. I spent a lot of time on this. It does not work. The model has to invent motion between two frames that were never meant to connect, and what it invents is a slow, strange morph. It looks worse than the jump.\n\nTwo honest choices:\n\nA half-screen is one frame, no first and no end. Attach the world for the light, the character or prop it shows, and write a small prompt. Name it after the scene it follows, with `-half`:\n\n```\n## 02 · half-screen, the eyes\n- Output: scenes/02-alarm-half.png\n- Attach: assets/room.png (light only) · assets/man.png\n\nSingle illustration, 16:9, the exact style of the attached images.\n[STYLE] man.png is the man: same face, hair and stubble.\n\nExtreme close-up of the man's face, filling the frame, on his side on\nthe pillow, eyes closed. Morning light across his face from the\nright. One eyebrow slightly raised, as if the ringing has just\nreached him. No bed edge, no clock, no room.\n```\n\n| After scene | Half-screen | Small motion in the clip | \n|---|---|---|\n| 01 · asleep | the clock face, close | the second hand ticks | \n| 02 · alarm | the man's closed eyes | the eyebrow lifts | \n| 03 · awake | bare feet touching the wooden floor | the toes curl | \n| 04 · walk | his hand on the white curtain | the curtain sways | \n| 05 · window | his eyes, open, with light in them | a slow blink | \n\nYou do not need all five. Use one where the jump is ugly, a plain cut where it is not.\n\nIn the edit, a half-screen **dissolves** in and out, half a second to a second on each side, over the scene before and the scene after. So a 4-second half-screen shows alone for two to three seconds. The eye is on the close-up while the room changes underneath it — that is what makes the jump disappear.\n\nA half-screen is a cut that looks like it was planned.\n\nYou may want to make the sound now — a voice from a voice model, the alarm, the city — and give it to the video model with the frames. You cannot. Not with the two frames.\n\nSeedance 2.5 runs on many platforms. I use it inside ElevenLabs — the same place I would make the voice — and even there, a voice file and a first-and-end frame pair cannot go into the same request. On fal.ai there is no audio input at all. On ByteDance's own API you can attach an audio reference, but the moment you do, the first and end frames stop being first and end — they become loose references, and the clip no longer runs from one to the other. I tried this more than once. Each time I got a good audio file and no place for it.\n\nSo the voice goes in the prompt. Write the line in quotes, describe the voice, say who speaks and when. The model renders the voice over the clip, with the mouth on the words, and makes the room sound too — the ring, the sheets, the far city. I have tested this many times. It is not perfect, but it is good, and it sits exactly where the picture needs it, because the same model made both.\n\n```\nHe stands at the window and says, in a low, tired voice, a man in\nhis thirties just awake: \"Morning.\" His mouth moves with the word.\nOnly he speaks.\nSound: his voice, the curtain, the city far below. No music.\n```\n\nHonest about quality: ElevenLabs' own voice model is better. But I cannot attach it next to my two frames, so it does not matter how good it is. Maybe a future version will take both. Until then, the best voice is the one you can actually put in the clip.\n\nThe voice you cannot attach is not a voice. It is a file.\n\nOne rule from here: 🎵 **no music in the clips.** Music goes over the whole film at the end, in one piece.\n\nOne clip per scene and per half-screen, into `video/`, prompts in `video.md`. For a scene, attach the first and the end frame. For a half-screen, the single frame.\n\nThree rules, and they all say **keep it short**:\n\n| Keep short | How | Why | \n|---|---|---|\n| ⏱️ **The clip** | 4 seconds | the model is at its best in short clips; long ones drift | \n| 🏃 **The motion** | one thing moves | the man walks, or the clock rings — not both | \n| ✍️ **The prompt** | a few lines | a long prompt makes worse motion, not better | \n\nWhy 4 seconds and not less? Because Seedance will not go lower. Many moments are shorter than that — a clock starts ringing in one second — but the clip is 4 seconds whether you need them or not. So put the motion in the **middle** of the clip and leave the first and last second quiet: still at the start, hold at the end. Those quiet seconds are what the edit fades over in Step 8. If the action starts on frame one, the fade eats it.\n\n```\n## 04 · The walk\n- Output: video/04-walk.mp4\n- Attach: scenes/04-walk-first.png (first) · scenes/04-walk-end.png (end) · 4 s\n\nImage-to-video, 4 seconds, from the first frame to the end frame.\nKeep the room, camera, man and bed exactly as drawn.\n0-1 s: he stands by the bed, still. 1-3 s: he walks slowly to the\nwindow, bare feet on wood. 3-4 s: he stops, back to us, his hand\nreaches for the curtain. Hold the end frame.\nSound: soft footsteps on wood. No music.\n## 04 · half-screen, the curtain\n- Output: video/04-walk-half.mp4\n- Attach: scenes/04-walk-half.png (first frame only) · 4 s\n\nImage-to-video, 4 seconds, from this single frame; the picture holds\nto the end. Small motion only: the curtain sways once, the fingers\ntighten on the cloth. No camera move.\nSound: the curtain, the city far away. No music.\n```\n\nThe temptation is to add the light, the mood, what the man feels. Every line you add, the model obeys by making the motion worse. Say what moves, say when, say what it sounds like, and stop.\n\nShort clip, one motion, few words. The frames do the talking.\n\nYou will remake some clips. When one is wrong, check the frames first — if the two frames do not agree, no prompt saves the clip. If the frames are right, cut the prompt, do not grow it.\n\nI use `ffmpeg`. It is free, it runs anywhere, and the edit becomes a text file you can read next month.\n\n📋 **The order.** One clip per line, and after each name, how that clip comes in: `cut` or `fade`. A scene after a scene is a cut. Anything touching a half-screen is a fade. This file *is* the edit. Save it as `film.txt`:\n\n```\n01-asleep        cut\n01-asleep-half   fade\n02-alarm         fade\n02-alarm-half    fade\n03-awake         fade\n03-awake-half    fade\n04-walk          fade\n04-walk-half     fade\n05-window        fade\n05-window-half   fade\n06-city          fade\n```\n\n🔗 **The join.** A `fade` overlaps two clips by `FADE` seconds, picture and sound together — one second is soft, half a second keeps more of the half-screen on screen. A `cut` is a one-frame crossfade: invisible, but it removes the click a hard cut leaves in the sound. This script reads the list and builds the chain:\n\n``` bash\n#!/bin/sh\n# cut.sh — joins the clips in film.txt into film.mp4\n# Every clip is LEN seconds. A fade overlaps two clips by FADE seconds.\nset -e\nLEN=4; FADE=1\nset --\nfor n in $(awk '{print $1}' film.txt); do set -- \"$@\" -i \"video/$n.mp4\"; done\nN=$(wc -l < film.txt | tr -d ' ')\nFILTER=$(awk -v len=\"$LEN\" -v fade=\"$FADE\" '\n  { n++; join[n] = $2 }\n  END {\n    v = \"[0:v]\"; a = \"[0:a]\"; t = len\n    for (i = 2; i <= n; i++) {\n      d = (join[i] == \"fade\") ? fade : 0.04\n      printf \"%s[%d:v]xfade=transition=fade:duration=%s:offset=%.2f[v%d];\", v, i-1, d, t - d, i\n      printf \"%s[%d:a]acrossfade=d=%s[a%d];\", a, i-1, d, i\n      v = \"[v\" i \"]\"; a = \"[a\" i \"]\"; t += len - d\n    }\n  }' film.txt)\nffmpeg -y \"$@\" -filter_complex \"${FILTER%;}\" -map \"[v$N]\" -map \"[a$N]\" \\\n  -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac film.mp4\n```\n\nRun `sh cut.sh`. Eleven clips of 4 seconds, ten fades of one second: **34 seconds** of film, no clicks. Change a line in `film.txt` — swap a `fade` for a `cut`, drop a half-screen — and run it again.\n\n🎵 **The finish.** A fade from black and to black, then the music over the whole film, quietly under the clips' own sound:\n\n```\nffmpeg -y -i film.mp4 \\\n  -vf \"fade=t=in:d=0.5,fade=t=out:st=33.5:d=0.5\" \\\n  -af \"afade=t=in:d=0.5,afade=t=out:st=33.5:d=0.5\" \\\n  film-faded.mp4\n\nffmpeg -y -i film-faded.mp4 -i music.mp3 \\\n  -filter_complex \"[1:a]volume=0.2,afade=t=out:st=30:d=4[m];[0:a][m]amix=inputs=2:duration=first\" \\\n  -c:v copy final.mp4\n```\n\nThe edit is a text file. Change one line, run it again.\n\nTwo models, and you can count every call before you start. List prices, October 2026: Seedance 2.5 at 720p with sound is $0.473 per second on fal.ai (other platforms differ). OpenAI bills GPT Image 2.5 by tokens and publishes no per-image price; the closest published figure is $0.165 for a 1536×1024 high-quality image. Run one call, read the `usage` field, and put your own number in.\n\n| Step | Model | What | Calls | Each | Cost | \n|---|---|---|---|---|---|\n| 2 🎨 | GPT Image 2.5 | 2 worlds, 1 character, 3 props | 6 | $0.17 | $0.99 | \n| 3 🎞️ | GPT Image 2.5 | 6 raw scenes | 6 | $0.17 | $0.99 | \n| 4 🔁 | GPT Image 2.5 | 6 second frames | 6 | $0.17 | $0.99 | \n| 5 👁️ | GPT Image 2.5 | 5 half-screens | 5 | $0.17 | $0.83 | \n| 7 🎥 | Seedance 2.5 | 11 clips × 4 s, at $0.473 / s | 11 | $1.89 | $20.81 | \n|  |  | **34 seconds of film** | **34** |  | **$24.61** | \n\nAbout twenty-five dollars, if every call comes out right the first time. It will not. Plan for every second call needing a retry, and the film costs closer to forty.\n\nAll twenty-three images together cost about as much as two clips. So the expensive mistake is never a bad image — it is a bad frame you only notice once it moves. That is why every step before 7 ends with \"do not move on until they match.\" Spend the cheap calls. Save the expensive ones.\n\nOne folder, five text files, and a film at the bottom. When something is wrong — and something will be — you know which file to open.\n\n`first` or `end`, and its partner is made from it\nIf you have been further than this — a model that takes a voice reference, a better way to bridge two scenes — tell me in the comments. I am still looking for the fix to the one step I could not make work.", "url": "https://wpnews.pro/news/ai-animation-from-idea-to-film-eight-small-jobs-instead-of-one-impossible-prompt", "canonical_source": "https://dev.to/dalirnet/ai-animation-from-idea-to-film-eight-small-jobs-instead-of-one-impossible-prompt-4p1b", "published_at": "2026-10-11 11:47:50+00:00", "updated_at": "2026-10-11 11:51:32.219894+00:00", "lang": "en", "topics": ["generative-ai", "ai-tools", "artificial-intelligence"], "entities": ["GPT Image 2.5 Sunburst"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-animation-from-idea-to-film-eight-small-jobs-instead-of-one-impossible-prompt", "markdown": "https://wpnews.pro/news/ai-animation-from-idea-to-film-eight-small-jobs-instead-of-one-impossible-prompt.md", "text": "https://wpnews.pro/news/ai-animation-from-idea-to-film-eight-small-jobs-instead-of-one-impossible-prompt.txt", "jsonld": "https://wpnews.pro/news/ai-animation-from-idea-to-film-eight-small-jobs-instead-of-one-impossible-prompt.jsonld"}}