cd /news/generative-ai/building-a-face-consistent-two-perso… · home › topics › generative-ai › article
[ARTICLE · art-146626] src=dev.to ↗ pub= topic=generative-ai verified=true sentiment=↑ positive

Building a Face-Consistent Two-Person Video Generator: Architecture, Models, and a Minimal Prototype

A developer published a build guide for a face-consistent two-person video generator, breaking the pipeline into five replaceable stages: face detection and alignment, identity embedding, motion and scene generation, audio, and compositing. The guide recommends insightface's buffalo_l model for detection and ArcFace-family 512-dim embeddings for identity preservation, and argues that face-swapping over a pre-rendered template is far cheaper than identity-conditioned video diffusion, which costs 10–100× the compute.

by read4 min views1 publishedOct 7, 2026

The "two photos in, one rap performance out" generators are everywhere right now. If you're a developer, the product is a curiosity — but the engineering underneath it is a genuinely interesting problem, and it maps onto a skill you'll reuse for a lot more than memes.

This post is a build guide, not an explainer. I'll walk through the architecture, the model choices at each stage, and a minimal prototype that gets you to "a face stays the same across a moving clip" without a PhD in diffusion.

Forget the orange stage. The core challenge in any personalized video generator is identity preservation: keep a specific face believable across dozens of frames while the body moves, the mouth animates, and the camera shifts.

A generic face can drift and nobody notices. Your user's face has to stay recognizably them for the whole clip, or the joke dies and the product looks broken. This is the axis everything else hangs on.

A working prototype is five stages. Each one is a replaceable component, which matters because you'll swap models at every layer.

upload ──► face detect/align ──► identity embed ──► motion+scene ──► audio ──► composite/export

You can't feed a raw selfie to a diffusion model. First you localize the face, crop it, and normalize it to a canonical orientation.

insightface) are the standard picks — fast on CPU, accurate enough.

from insightface.app import FaceAnalysis

app = FaceAnalysis(name="buffalo_l", providers=["CPUExecutionProvider"])
app.prepare(ctx_id=0, det_size=(640, 640))

img = load_image("person_a.jpg")
faces = app.get(img)          # each face has .bbox, .kps, .normed_embedding

The normed_embedding on each detected face is the prize — it's the identity vector you'll use downstream.

This is the step most tutorials skip. You need a compact vector that represents the person, independent of pose, lighting, and expression. Face recognition models trained on large datasets (ArcFace-family backbones) give you exactly that — a 512-dim embedding where "same person" = "high cosine similarity."

Keep the embedding for the source face (the user's photo). In generation, you'll pull the generated face back toward this vector.

Here you have two real architectural choices, and they're not equivalent:

Approach How it works Trade-off
Face-swap over a template Generate or reuse a fixed performance clip, then swap in the target face frame-by-frame Fast, cheap, identity locked by construction; less control over scene
Identity-conditioned video diffusion Drive a video model with the identity embedding (IP-Adapter-style injection or LoRA) More flexible scenes; harder to keep identity stable, higher compute

For a two-person rap clip, the first approach is why these tools feel instant: the expensive part (a believable performance with choreography and camera) is solved once, at build time, and reused for every user. You're not generating a scene from scratch per request — you're compositing faces into a pre-rendered one. That's a cost and latency decision, not just a convenience.

If you want bespoke scenes, look at identity-injection via IP-Adapter + AnimateDiff, or a fine-tuned LoRA per recurring character. Budget 10–100× the compute.

Don't synthesize the copyrighted reference track. Generate an original beat plus lyrics and vocals — typically a music/lyrics LLM for the bars, then a TTS or singing-vocals model (Bark, or a commercial text-to-song API) for the performance. Keep audio as a separate pass and mux it at the end; coupling it to the video model only adds failure modes.

Stitch frames with ffmpeg, mux the audio, encode for social (vertical, H.264). Nothing exotic here, but get the frame rate and aspect ratio right before you scale.

If you want something working in a weekend, the face-swap route is the pragmatic choice. The modern standard is insightface's inswapper model:

import cv2, numpy as np
from insightface.app import FaceAnalysis
import onnxruntime as ort

swapper = insightface.model_zoo.get_model("inswapper_128.onnx",
                                          providers=["CPUExecutionProvider"])
source = app.get(load_image("person_a.jpg"))[0]      # identity to inject

video = cv2.VideoCapture("template_performance.mp4")
writer = setup_writer("out.mp4", fps=video.get(cv2.CAP_PROP_FPS))

while True:
    ok, frame = video.read()
    if not ok:
        break
    targets = app.get(frame)
    for t in targets:
        frame = swapper.get(frame, t, source, paste_back=True)
    writer.write(frame)

writer.release()

This gives you a template performance with a stable face in under a hundred lines. Two people? Run the swap for each of two source identities, or use a template that already has two performers in fixed left/right positions and assign each face to its slot.

The prototype gets you 80% there. The last 20% is where products are won or lost:

If this is your product, build the pipeline — the identity-preservation problem is a genuine moat worth owning.

If you just need the output — say, a shareable two-person clip for a campaign or a quick test of the format — don't re-solve a solved problem. Something like AI Rap Duo already bundles the template, the identity preservation, and the original beat into a two-photo upload with no prompt. That's the right call when the goal is the video, not the architecture.

Either way, the interesting engineering question is the same one that makes any personalized AI product hard: keeping a real, specific face believable while everything around it moves. Nail that, and the rest is just plumbing.

── more in #generative-ai 4 stories · sorted by recency
── more on @insightface 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/building-a-face-cons…] indexed:0 read:4min 2026-10-07 · —