# Building a Face-Consistent Two-Person Video Generator: Architecture, Models, and a Minimal Prototype

> Source: <https://dev.to/marita_pang_0302f784f7ea3/building-a-face-consistent-two-person-video-generator-architecture-models-and-a-minimal-prototype-k5l>
> Published: 2026-10-07 06:10:42+00:00

The "two photos in, one rap performance out" generators are everywhere right now. If you're a developer, the product is a curiosity — but the engineering underneath it is a genuinely interesting problem, and it maps onto a skill you'll reuse for a lot more than memes.

This post is a build guide, not an explainer. I'll walk through the architecture, the model choices at each stage, and a minimal prototype that gets you to "a face stays the same across a moving clip" without a PhD in diffusion.

Forget the orange stage. The core challenge in any personalized video generator is **identity preservation**: keep *a specific face* believable across dozens of frames while the body moves, the mouth animates, and the camera shifts.

A generic face can drift and nobody notices. Your user's face has to stay recognizably *them* for the whole clip, or the joke dies and the product looks broken. This is the axis everything else hangs on.

A working prototype is five stages. Each one is a replaceable component, which matters because you'll swap models at every layer.

```
upload ──► face detect/align ──► identity embed ──► motion+scene ──► audio ──► composite/export
```

You can't feed a raw selfie to a diffusion model. First you localize the face, crop it, and normalize it to a canonical orientation.

`insightface`) are the standard picks — fast on CPU, accurate enough.

``` python
from insightface.app import FaceAnalysis

app = FaceAnalysis(name="buffalo_l", providers=["CPUExecutionProvider"])
app.prepare(ctx_id=0, det_size=(640, 640))

img = load_image("person_a.jpg")
faces = app.get(img)          # each face has .bbox, .kps, .normed_embedding
```

The `normed_embedding` on each detected face is the prize — it's the identity vector you'll use downstream.

This is the step most tutorials skip. You need a compact vector that *represents* the person, independent of pose, lighting, and expression. Face recognition models trained on large datasets (ArcFace-family backbones) give you exactly that — a 512-dim embedding where "same person" = "high cosine similarity."

Keep the embedding for the *source* face (the user's photo). In generation, you'll pull the generated face back toward this vector.

Here you have two real architectural choices, and they're not equivalent:

| Approach | How it works | Trade-off | 
|---|---|---|
| **Face-swap over a template** | Generate or reuse a fixed performance clip, then swap in the target face frame-by-frame | Fast, cheap, identity locked by construction; less control over scene | 
| **Identity-conditioned video diffusion** | Drive a video model with the identity embedding (IP-Adapter-style injection or LoRA) | More flexible scenes; harder to keep identity stable, higher compute | 

For a two-person rap clip, the first approach is why these tools feel instant: the expensive part (a believable performance with choreography and camera) is solved once, at build time, and reused for every user. You're not generating a scene from scratch per request — you're compositing faces into a pre-rendered one. That's a cost and latency decision, not just a convenience.

If you want bespoke scenes, look at identity-injection via IP-Adapter + AnimateDiff, or a fine-tuned LoRA per recurring character. Budget 10–100× the compute.

Don't synthesize the copyrighted reference track. Generate an original beat plus lyrics and vocals — typically a music/lyrics LLM for the bars, then a TTS or singing-vocals model (Bark, or a commercial text-to-song API) for the performance. Keep audio as a separate pass and mux it at the end; coupling it to the video model only adds failure modes.

Stitch frames with `ffmpeg`, mux the audio, encode for social (vertical, H.264). Nothing exotic here, but get the frame rate and aspect ratio right before you scale.

If you want something working in a weekend, the face-swap route is the pragmatic choice. The modern standard is `insightface`'s inswapper model:

``` python
import cv2, numpy as np
from insightface.app import FaceAnalysis
import onnxruntime as ort

# Load models once
swapper = insightface.model_zoo.get_model("inswapper_128.onnx",
                                          providers=["CPUExecutionProvider"])
source = app.get(load_image("person_a.jpg"))[0]      # identity to inject

video = cv2.VideoCapture("template_performance.mp4")
writer = setup_writer("out.mp4", fps=video.get(cv2.CAP_PROP_FPS))

while True:
    ok, frame = video.read()
    if not ok:
        break
    targets = app.get(frame)
    for t in targets:
        frame = swapper.get(frame, t, source, paste_back=True)
    writer.write(frame)

writer.release()
```

This gives you a template performance with a stable face in under a hundred lines. Two people? Run the swap for each of two source identities, or use a template that already has two performers in fixed left/right positions and assign each face to its slot.

The prototype gets you 80% there. The last 20% is where products are won or lost:

If this is your product, build the pipeline — the identity-preservation problem is a genuine moat worth owning.

If you just need the *output* — say, a shareable two-person clip for a campaign or a quick test of the format — don't re-solve a solved problem. Something like [AI Rap Duo](https://airapduo.com/) already bundles the template, the identity preservation, and the original beat into a two-photo upload with no prompt. That's the right call when the goal is the video, not the architecture.

Either way, the interesting engineering question is the same one that makes *any* personalized AI product hard: keeping a real, specific face believable while everything around it moves. Nail that, and the rest is just plumbing.
