{"slug": "5-things-i-learned-from-bad-ai-video-generations", "title": "5 Things I Learned From Bad AI Video Generations", "summary": "A developer who tested AI video generation models found that vague motion descriptors like \"dynamic tracking shot\" often fail to produce the intended camera movement, and that explicitly describing camera paths, subject action, and environmental reactions yields more accurate results. The engineer documented five lessons from failed generations, including separating subject motion from camera motion and specifying how movement affects the surrounding scene.", "body_md": "AI video prompts can look completely reasonable and still produce a result that is not what you expected.\n\nI ran into this while testing a skiing video.\n\nThe skier moved downhill correctly, the snow reacted to the skis, and the overall scene looked good. The problem was the camera. I wanted it to move in behind the skier, come closer, pass near the subject, and then pull away.\n\nInstead, it barely moved.\n\nMy prompt included the phrase *\"dynamic tracking shot,\"* so at first I thought the model had simply ignored the camera instruction. But when I looked at the prompt again, I realized that I had not actually described the camera movement in much detail.\n\nThat test changed the way I write AI video prompts.\n\nI now spend less time adding descriptive words and more time separating the different kinds of movement inside the shot.\n\nHere are five things that have helped me.\n\nOne problem with AI video prompts is that they can sound like video prompts while still mostly describing a static image.\n\nFor example:\n\n*A skier on a snowy mountain, cinematic lighting, dramatic winter landscape, dynamic movement.*\n\nThere is nothing technically wrong with this prompt, but most of it describes the appearance of the scene.\n\nThe model still has to decide what *dynamic movement* means.\n\nI usually get a more useful result when I describe the action directly.\n\nInstead of:\n\n*A skier moving through the snow.*\n\nI would write:\n\n*The skier accelerates downhill from left to right, leans into a sharp turn, and cuts across the slope.*\n\nThe second version gives the model more information about direction and sequence.\n\nThis does not mean every movement needs to be described in extreme detail. Long prompts can create their own problems. I mainly look for vague verbs or adjectives that are standing in for an action I could describe more clearly.\n\nWords such as *dynamic*, *cinematic*, or *energetic* can help describe the overall feeling of a shot, but they usually do not replace an actual motion instruction.\n\nThis was the main problem in my skiing test.\n\nThe subject motion was already acceptable. The skier moved naturally, and the skis kicked up snow during the turn.\n\nThe camera was the part that did not match what I had imagined.\n\nMy original instruction was simply:\n\n*dynamic tracking shot*\n\nI knew what I meant by that. The model did not necessarily know which version of a tracking shot I wanted.\n\nSo I rewrote the camera instruction as a path:\n\n*The camera approaches from behind the skier, moves alongside them, closes the distance, passes near the subject, then separates as snow bursts across the frame.*\n\nThe result was much closer to the movement I wanted.\n\nSince then, I still use terms such as *tracking shot*, *dolly shot*, or *aerial shot*, but I do not rely on them when the exact camera movement matters.\n\nFor example, there is a fairly big difference between:\n\n*The camera tracks the runner.*\n\nand:\n\n*The camera begins behind the runner, moves toward their left side, briefly matches their speed, then pulls ahead.*\n\nBoth can describe a tracking shot, but the second one gives the model an actual route through the scene.\n\nSometimes the main subject is moving correctly and the video still feels strangely static.\n\nIn those cases, I check what is happening around the subject.\n\nIn the skiing example, this made a noticeable difference.\n\nCompare:\n\n*The skier makes a sharp turn.*\n\nwith:\n\n*The skier makes a sharp turn. The skis cut into the snow and throw a burst of powder outward as the skier passes.*\n\nThe second prompt is not only describing the skier. It is also describing what the movement does to the environment.\n\nThis can apply to many types of shots.\n\nA person running may move their clothes or hair. A car may throw up dust or spray water from the road. Someone walking through curtains should probably move the fabric. An object falling into water should affect the surface around it.\n\nI do not add environmental reactions to every prompt because that can make a simple scene unnecessarily complicated.\n\nBut when the subject is doing the correct action and the shot still feels flat, this is one of the things I check.\n\nI noticed a different problem while testing a recurring character called Carrie.\n\nI had several reference images for her, including front-facing, profile, and full-body views.\n\nAt first, I still wrote long descriptions of her appearance inside the video prompt:\n\n*A young woman with [face description], [hair description], [clothing description]...*\n\nAfter testing the workflow a few times, I found that much of this information was already coming from the reference image.\n\nThat meant the prompt could focus more on the action.\n\n*Carrie walks toward the table, pulls out the chair, sits down, and looks toward the window.*\n\nThe reference image handles the character identity, while the prompt describes what happens in the shot.\n\nI also found that the type of reference image matters.\n\nFor a close-up, a clear front-facing reference can be useful. For a walking sequence or another shot where body proportions matter, I prefer a wider or full-body reference.\n\nThis also makes failed generations easier to troubleshoot.\n\nIf the character identity changes too much, I check the reference setup first.\n\nIf the character still looks correct but performs the wrong movement, I look at the motion prompt.\n\nBefore I started separating those two problems, my instinct was often to keep adding more description to the same prompt. That usually made it harder to tell which instruction was actually helping.\n\nI learned this while making a 15-second bookstore sequence.\n\nThe basic idea was simple: a character enters a bookstore, walks between the shelves, finds a book, opens it, and ends in a quieter final shot.\n\nWhen I wrote the whole sequence as one paragraph, it became difficult to see whether the timing and progression actually made sense.\n\nSo I planned it as separate beats first:\n\n**Shot 1:** Establish the bookstore.\n\n**Shot 2:** Follow the character between the shelves.\n\n**Shot 3:** Move closer as they reach for a book.\n\n**Shot 4:** Show them opening it.\n\n**Shot 5:** End on a quieter frame.\n\nThis does not mean I always generate five separate clips.\n\nIn this test, generating the shots independently sometimes caused a different problem. The position of the shelves, the character, and the camera changed between clips, so the sequence no longer felt like the same physical space.\n\nA longer continuous generation could keep the environment more coherent.\n\nThe shot list was still useful, though, because it gave me a way to plan the timing and progression before deciding how many generations to make.\n\nI now treat shot planning and clip generation as two separate decisions.\n\nA video can be planned as five shots and still be generated as one continuous sequence.\n\nAfter these tests, I started using four categories when a motion prompt is not working:\n\n**Subject motion + Camera path + Environmental reaction + Temporal sequence**\n\nFor the skiing video, I might break the prompt down like this:\n\n**Subject motion**\n\n*The skier accelerates downhill and leans into a sharp turn.*\n\n**Camera path**\n\n*The camera approaches from behind, moves alongside the skier, then closes the distance as they pass.*\n\n**Environmental reaction**\n\n*The skis cut into the snow and throw powder toward the camera.*\n\n**Temporal sequence**\n\n*The shot begins wide, moves into a close tracking moment, then pulls away as the skier continues downhill.*\n\nI do not necessarily write every prompt in this exact format.\n\nIt is more useful as a troubleshooting checklist.\n\nIf the skier is moving correctly but the camera is static, I probably do not need to rewrite the subject description.\n\nIf the camera moves correctly but the scene still feels lifeless, I can look at environmental reactions.\n\nIf all the individual actions are there but they happen in the wrong order, the temporal sequence may need to be clearer.\n\nAI video generation is still unpredictable. A model can ignore an instruction or introduce something that was never in the prompt.\n\nBut breaking motion down this way has made failed generations easier for me to diagnose.\n\nAnd in practice, that has been more useful than simply making the prompt longer.\n\nWhat usually goes wrong first in your AI video generations: subject motion, camera movement, character consistency, or timing?\n\n#ai #promptengineering #videogeneration #tutorial", "url": "https://wpnews.pro/news/5-things-i-learned-from-bad-ai-video-generations", "canonical_source": "https://dev.to/lee_xiaoyuan_a97212d2f33b/5-things-i-learned-from-bad-ai-video-generations-3d58", "published_at": "2026-09-23 00:44:20+00:00", "updated_at": "2026-09-23 01:22:43.409326+00:00", "lang": "en", "topics": ["generative-ai", "ai-tools", "ai-products"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/5-things-i-learned-from-bad-ai-video-generations", "markdown": "https://wpnews.pro/news/5-things-i-learned-from-bad-ai-video-generations.md", "text": "https://wpnews.pro/news/5-things-i-learned-from-bad-ai-video-generations.txt", "jsonld": "https://wpnews.pro/news/5-things-i-learned-from-bad-ai-video-generations.jsonld"}}