How to Evaluate an AI Video Character Replacement A developer working on Genjutsu AI documented a case study of an AI video character-replacement test — a 1280×720, roughly 8.04-second talking-shot clip — and outlined a review process for evaluating such outputs. The recorded review found the appearance and clothing changed while the background was broadly retained and the original audio stayed aligned, but the developer cautions that a single successful clip establishes no success rate and does not prove face consistency, natural hands, or accurate lip sync. The proposed process separates audio alignment from lip movement, samples the beginning, middle, and end of a clip, and applies identical criteria across versions to keep conclusions traceable to the footage and settings. An AI video edit can look convincing in one frame and fall apart a second later. For character replacement, the useful question is whether the new appearance holds up while the original performance continues. Here is a small case study from a project I work on, followed by a review process developers can use when evaluating similar outputs. The documented test used: The resulting file was 1280 × 720 and approximately 8.04 seconds long. Keeping these details matters. A different model, reference, or preprocessing step introduces another variable. If you change several inputs together, it becomes difficult to explain why the result improved. A useful test record includes the source segment, reference image, prompt, model, settings, output file, and review notes. The recorded review reported changed appearance and clothing, a broadly retained background at the start, middle, and end, and aligned original audio. Those observations support a limited conclusion: this particular talking-shot example preserved several intended properties during a basic review. They do not establish precise face consistency, natural hands, or accurate lip sync throughout every frame. They also do not establish reliable results for dancing, heavy occlusion, or replacing one person while leaving another entirely unchanged. A single successful clip helps define the next test. It does not provide a success rate. For a character-replacement task, define the expected changes and the properties that should remain stable. A practical review covers: Audio alignment and lip movement deserve separate checks. Retaining the original soundtrack says little about whether the generated mouth moves convincingly. Compare matching moments in the source and output. Sample the beginning, middle, and end, then inspect transitions and obstructed frames more closely. Finally, watch the entire clip at normal speed with sound. If the identity drifts, a clearer reference photo or a simpler source segment is a useful next experiment. If the scene changes unexpectedly, check the selected task and preservation instructions. A motion-transfer task may reconstruct the surroundings rather than preserve them. Use the same review criteria for each version. Otherwise, it is easy to favor the output with the strongest opening frame and overlook a new problem later in the clip. For a broader evaluation, add cases with side profiles, fast gestures, occlusions, and multiple people. Track failures as carefully as successful examples. The practical goal is to make each conclusion traceable to the footage and settings behind it. That makes both product decisions and claims about quality easier to defend. Disclosure: The documented example comes from Genjutsu AI, a project I work on. This article was drafted with AI using the project's documented test material.