I Built an AI Text-to-Video Generator — Here's What I Learned About Prompt Engineering for Video A developer who built the AI text-to-video generator CineGen shared lessons from hundreds of test renders on how video prompt engineering differs from image prompting. The core finding is that video prompts must specify motion and camera behavior — what is in frame, what is moving, and how the shot evolves — rather than relying on static scene descriptions and adjectives, with cinematography terms like 'tracking shot' and 'slow dolly in' constraining the motion field and reducing artifacts. When I started building CineGen https://www.cine-gen.com , an AI text-to-video generator, I assumed prompt engineering for video would be "image prompting, plus the word moving ." It is not. After hundreds of test renders and a lot of embarrassing outputs a seagull with nine wings remains burned in my memory , I learned that prompting for video is closer to directing a 5-second film than describing a photograph. Here are the lessons that actually changed the quality of my outputs. No hype, just what works. The single biggest mindset shift: an image prompt describes one moment . A video prompt describes a sequence of moments . If your prompt only describes a static scene, the model has to invent the motion — and it will invent something weird. Before image thinking : a beautiful sunset over the ocean, cinematic After video thinking : Wide aerial shot of an ocean at sunset. The camera slowly pans left across the water as waves roll toward the shore. Orange light flickers on the wave crests. A small sailboat crosses the frame from right to left. Cinematic lighting, calm mood. The second prompt works because it answers three questions the model needs: what's in the frame, what's moving, and how does the shot evolve? A structure I keep coming back to is the three-beat arc: establish → action → resolve . Even in a 4-second clip, giving the model a beginning, middle, and end dramatically reduces the "slideshow of random frames" effect. This was the highest-leverage discovery. Generative video models respond strongly to cinematography vocabulary — much more than you'd expect. Naming the shot is often more effective than describing the scene in detail. A short glossary that covers ~90% of what I use: extreme close-up , close-up , medium shot , wide shot , aerial shot slow dolly in , pan left/right , tilt up/down , tracking shot , static shot , orbit around shallow depth of field , 35mm , handheld , smooth gimbal motion Before: a person walking through a forest After: Tracking shot following a hiker from behind on a forest trail, camera gliding smoothly at walking pace. Tall pine trees blur past on both sides, morning fog drifting between trunks. Shallow depth of field, natural light. The difference is night and day. "Tracking shot" tells the model how the camera behaves, which constrains the motion field and kills a huge class of artifacts where the background slides around unnaturally. If you take one thing from this article: direct the camera, not just the scene. One caution: don't stack contradictory camera moves. dolly in + pan left + tilt up + orbit in one prompt is asking the model to solve an impossible motion puzzle. One primary camera move per clip. In image prompting, adjectives carry the load: beautiful, stunning, ultra-detailed . In video prompting, most adjectives are noise. What the model needs is motion specification — verbs with direction, speed, and rhythm. Compare: Adjective-heavy weak a stunning beautiful waterfall in a gorgeous lush forest, amazing Verb-heavy strong Waterfall plunging down a mossy cliff into a pool below, mist rising and drifting left. Ferns swaying gently in the foreground. Camera holds a static wide shot. Notice the second prompt barely uses adjectives, yet produces a far better clip. My rule of thumb: every noun in the prompt should have a verb attached to it. If something is in the frame, say what it's doing — even if it's just "standing still" which, by the way, is a legitimate and useful instruction: the cat sits perfectly still, only its tail flicks . Speed words matter too: slowly , gently , rapidly , suddenly . Models genuinely differentiate these. "Walks slowly toward the camera" and "runs toward the camera" produce very different motion — use that dial deliberately. The hardest problem in AI video isn't making pretty frames — it's making frame 1 and frame 48 agree with each other. Faces morph, jackets change color, a coffee cup teleports between hands. Here's what actually helps: Anchor the subject with specific, repeated attributes. Don't write "a woman"; write "a woman with short black hair in a red jacket." The more specific the anchor, the harder it is for the model to drift. Color anchors red jacket work especially well because color is one of the more stable features across frames. One action per clip. This is the constraint I resisted longest and benefited from most. "She picks up the cup, drinks, sets it down, and waves" will break. "She lifts the cup and takes a sip" works. Complex multi-stage actions across a few seconds are where morphing artifacts breed. Chain short clips instead of cramming everything into one prompt. Avoid mid-scene transformations. Prompts like "the car transforms into a robot" or "day turns to night" ask the model to do the single hardest thing in generative video: coherent metamorphosis. It will produce something , but it won't be what you pictured. Keep state changes out of the prompt; do them as separate clips and cut between them. In image generation, negative prompts are optional polish. In video, they're load-bearing, because video has failure modes that images don't: flickering, morphing, limb duplication during motion, warping geometry. A negative prompt block I reuse constantly: negative prompt: morphing face, extra limbs, extra fingers, flickering, warping background, distorted hands, text, watermark, sudden scene change, deformed body, disappearing objects A few notes on this list: flickering and warping background target specifically temporal artifacts — these do almost nothing for still images but matter enormously for video. text, watermark — generated text in video is doubly cursed: not only is it usually gibberish, it sudden scene change suppresses the model's urge to "cut" mid-clip when it gets confused, which reads as a glitch rather than an edit. Keep the negative list focused. A 40-item negative prompt dilutes into noise; 8–12 targeted terms beat a kitchen sink. After all this trial and error, I converged on a template. It's boring, and that's the point — boring is reproducible: SHOT + SUBJECT + ANCHORS + ACTION with verbs + ENVIRONMENT + LIGHTING / MOOD + CAMERA MOVEMENT + STYLE TAG Filled in: Medium shot of a barista with tied-back brown hair and a denim apron, pouring steamed milk into a ceramic cup to form latte art. Warm morning light through a cafe window, dust motes drifting in the sunbeam. Camera slowly dollies in. Photorealistic, shallow depth of field. Negative: morphing face, extra fingers, flickering, warping background, text, watermark Every slot filled, one camera move, one action, anchored subject, targeted negatives. This structure won't win avant-garde awards, but it produces usable clips at a dramatically higher hit rate than freeform prose. When a render fails, the template also makes debugging easy: you know exactly which slot to change. Some things just don't work yet, regardless of prompt craft. Knowing the walls saves you render credits and frustration: My workaround for all of these is the same: compose around the weakness. Can't do hands? Frame the shot so hands are out of view. Can't do text? Make the sign out of focus. Prompt engineering isn't just writing better prompts — it's choosing shots the model can actually execute. The workflow that finally made this sustainable: The unglamorous truth: great AI video comes from iteration discipline, not from one magical prompt. The people getting the best results aren't better writers — they're better experimenters. That's the honest version of what I learned building CineGen https://www.cine-gen.com . Video prompting rewards directors, not poets: think in timelines, speak in camera moves, anchor everything, and iterate like an engineer. If you want to put this into practice without wrestling with local GPU setups, CineGen is the text-to-video tool I built around exactly this workflow — describe your shot, iterate fast, and keep the clips that work. There's a Pro plan at $9.90/month and a $199 lifetime deal if you'd rather not do subscriptions. Either way, go make something weird — the nine-winged seagull era of your prompting journey is waiting.