Building a Face-Consistent Two-Person Video Generator: Architecture, Models, and a Minimal Prototype A developer published a build guide for a face-consistent two-person video generator, breaking the pipeline into five replaceable stages: face detection and alignment, identity embedding, motion and scene generation, audio, and compositing. The guide recommends insightface's buffalo_l model for detection and ArcFace-family 512-dim embeddings for identity preservation, and argues that face-swapping over a pre-rendered template is far cheaper than identity-conditioned video diffusion, which costs 10–100× the compute. The "two photos in, one rap performance out" generators are everywhere right now. If you're a developer, the product is a curiosity — but the engineering underneath it is a genuinely interesting problem, and it maps onto a skill you'll reuse for a lot more than memes. This post is a build guide, not an explainer. I'll walk through the architecture, the model choices at each stage, and a minimal prototype that gets you to "a face stays the same across a moving clip" without a PhD in diffusion. Forget the orange stage. The core challenge in any personalized video generator is identity preservation : keep a specific face believable across dozens of frames while the body moves, the mouth animates, and the camera shifts. A generic face can drift and nobody notices. Your user's face has to stay recognizably them for the whole clip, or the joke dies and the product looks broken. This is the axis everything else hangs on. A working prototype is five stages. Each one is a replaceable component, which matters because you'll swap models at every layer. upload ──► face detect/align ──► identity embed ──► motion+scene ──► audio ──► composite/export You can't feed a raw selfie to a diffusion model. First you localize the face, crop it, and normalize it to a canonical orientation. insightface are the standard picks — fast on CPU, accurate enough. python from insightface.app import FaceAnalysis app = FaceAnalysis name="buffalo l", providers= "CPUExecutionProvider" app.prepare ctx id=0, det size= 640, 640 img = load image "person a.jpg" faces = app.get img each face has .bbox, .kps, .normed embedding The normed embedding on each detected face is the prize — it's the identity vector you'll use downstream. This is the step most tutorials skip. You need a compact vector that represents the person, independent of pose, lighting, and expression. Face recognition models trained on large datasets ArcFace-family backbones give you exactly that — a 512-dim embedding where "same person" = "high cosine similarity." Keep the embedding for the source face the user's photo . In generation, you'll pull the generated face back toward this vector. Here you have two real architectural choices, and they're not equivalent: | Approach | How it works | Trade-off | |---|---|---| | Face-swap over a template | Generate or reuse a fixed performance clip, then swap in the target face frame-by-frame | Fast, cheap, identity locked by construction; less control over scene | | Identity-conditioned video diffusion | Drive a video model with the identity embedding IP-Adapter-style injection or LoRA | More flexible scenes; harder to keep identity stable, higher compute | For a two-person rap clip, the first approach is why these tools feel instant: the expensive part a believable performance with choreography and camera is solved once, at build time, and reused for every user. You're not generating a scene from scratch per request — you're compositing faces into a pre-rendered one. That's a cost and latency decision, not just a convenience. If you want bespoke scenes, look at identity-injection via IP-Adapter + AnimateDiff, or a fine-tuned LoRA per recurring character. Budget 10–100× the compute. Don't synthesize the copyrighted reference track. Generate an original beat plus lyrics and vocals — typically a music/lyrics LLM for the bars, then a TTS or singing-vocals model Bark, or a commercial text-to-song API for the performance. Keep audio as a separate pass and mux it at the end; coupling it to the video model only adds failure modes. Stitch frames with ffmpeg , mux the audio, encode for social vertical, H.264 . Nothing exotic here, but get the frame rate and aspect ratio right before you scale. If you want something working in a weekend, the face-swap route is the pragmatic choice. The modern standard is insightface 's inswapper model: python import cv2, numpy as np from insightface.app import FaceAnalysis import onnxruntime as ort Load models once swapper = insightface.model zoo.get model "inswapper 128.onnx", providers= "CPUExecutionProvider" source = app.get load image "person a.jpg" 0 identity to inject video = cv2.VideoCapture "template performance.mp4" writer = setup writer "out.mp4", fps=video.get cv2.CAP PROP FPS while True: ok, frame = video.read if not ok: break targets = app.get frame for t in targets: frame = swapper.get frame, t, source, paste back=True writer.write frame writer.release This gives you a template performance with a stable face in under a hundred lines. Two people? Run the swap for each of two source identities, or use a template that already has two performers in fixed left/right positions and assign each face to its slot. The prototype gets you 80% there. The last 20% is where products are won or lost: If this is your product, build the pipeline — the identity-preservation problem is a genuine moat worth owning. If you just need the output — say, a shareable two-person clip for a campaign or a quick test of the format — don't re-solve a solved problem. Something like AI Rap Duo https://airapduo.com/ already bundles the template, the identity preservation, and the original beat into a two-photo upload with no prompt. That's the right call when the goal is the video, not the architecture. Either way, the interesting engineering question is the same one that makes any personalized AI product hard: keeping a real, specific face believable while everything around it moves. Nail that, and the rest is just plumbing.