# Build a three-stage video dedup worker: SHA-256, perceptual hash, then FFmpeg's MPEG-7 signature

> Source: <https://dev.to/masonwritescode/build-a-three-stage-video-dedup-worker-sha-256-perceptual-hash-then-ffmpegs-mpeg-7-signature-578l>
> Published: 2026-09-29 09:12:35+00:00

## TL;DR

We'll build a dedup worker that catches re-encoded, resized, watermarked and trimmed copies of videos you already have. Stage 1 is a byte hash, stage 2 is a 64-bit perceptual hash checked by Hamming distance, stage 3 is FFmpeg's built-in MPEG-7 `signature` filter for "does A contain B, and where". Python 3.12, FFmpeg 7 or newer, SQLite.

A `dedup.py` module with one entry point, `check(path) -> Verdict`, that returns `exact`, `near`, `contains`, or `new`, plus a SQLite index so it gets faster as the library grows. No ML, no GPUs, nothing that isn't `pip install` or `apt install`.

The reason for three stages instead of one: the cheap check should run on every upload before you spend money transcoding, and the expensive check should only run on candidates the cheap check surfaced.

```
sudo apt-get install -y ffmpeg     # 7.x or 8.x; the signature filter has shipped since 2017
python3 -m venv .venv && source .venv/bin/activate
pip install videohash2            # maintained fork of videohash; pin the current release in requirements.txt
ffmpeg -hide_banner -filters | grep signature
```

You should see:

``` php
 ... signature          N->V       Calculate the MPEG-7 video signature
```

If that line is missing, your FFmpeg build was configured without it; grab a static build.

We'll synthesise a source and three "duplicates" so the tests are reproducible and the repo stays binary-free.

```
# scripts/make_fixtures.sh
set -e
mkdir -p fixtures && cd fixtures

# 20-second source
ffmpeg -y -f lavfi -i "testsrc2=duration=20:size=640x360:rate=25,format=yuv420p" \
       -f lavfi -i "sine=frequency=440:duration=20" \
       -c:v libx264 -crf 20 -c:a aac -shortest source.mp4

# A: re-encode at lower quality and resolution
ffmpeg -y -i source.mp4 -vf scale=320:180 -c:v libx264 -crf 32 -c:a aac dup_reencode.mp4

# B: watermark in the corner
ffmpeg -y -i source.mp4 -vf "drawbox=x=10:y=10:w=80:h=30:color=white@0.8:t=fill" \
       -c:v libx264 -crf 23 -c:a copy dup_watermark.mp4

# C: 6-second excerpt starting at 0:07 (containment case)
ffmpeg -y -ss 7 -i source.mp4 -t 6 -c:v libx264 -crf 23 -c:a aac excerpt.mp4

# D: unrelated video
ffmpeg -y -f lavfi -i "mandelbrot=size=640x360:rate=25" -t 20 \
       -c:v libx264 -crf 23 unrelated.mp4
```

`sha256sum fixtures/*.mp4` gives you four different hashes for what a human would call two videos. That's the problem.

``` python
# dedup/stage1.py
import hashlib
from pathlib import Path

def sha256(path: Path, chunk: int = 1 << 20) -> str:
    h = hashlib.sha256()
    with path.open("rb") as f:
        while blk := f.read(chunk):
            h.update(blk)
    return h.hexdigest()
```

This catches the accidental double upload and nothing else. It stays because it costs nothing and it's the only stage with zero false positives.

`videohash2` samples one frame per second, shrinks each to 144x144, tiles them into a collage, and wavelet-hashes the collage into 64 bits. Two re-encodes of the same clip land a few bits apart.

``` python
# dedup/stage2.py
from pathlib import Path
from videohash2 import VideoHash

def phash(path: Path) -> int:
    vh = VideoHash(path=str(path))
    return int(vh.hash_hex, 16)          # 64-bit int

def hamming(a: int, b: int) -> int:
    return (a ^ b).bit_count()
```

Try it on the fixtures:

``` python
>>> from pathlib import Path
>>> from dedup.stage2 import phash, hamming
>>> src = phash(Path("fixtures/source.mp4"))
>>> for name in ["dup_reencode", "dup_watermark", "excerpt", "unrelated"]:
...     print(name, hamming(src, phash(Path(f"fixtures/{name}.mp4"))))
```

You'll see the re-encode and the watermark land close to zero and `unrelated` land far away. The interesting one is `excerpt`: it will **not** be close. That's by design and the library says so in its docs: a whole-video collage hash cannot tell you that one video is a part of another. That's what stage 3 is for.

⚠️ Note: pick the Hamming threshold conservatively and log near-misses. A false positive here means one user's upload gets served as another user's asset. Start tight (single digits), review what you're rejecting, and loosen only with evidence.

FFmpeg's `signature` filter fingerprints every frame and, given two or more inputs, searches for matching sequences. First, generate and **store** a signature per asset. Generation is the slow part; the maintainer of one tool built on this filter reports 19 minutes to fingerprint 96 files versus 45 seconds to compare them all. Never regenerate.

```
# one-time per asset: write a binary signature next to it
ffmpeg -hide_banner -i fixtures/source.mp4 \
  -vf "signature=format=binary:filename=fixtures/source.sig" -f null -
```

Then compare two inputs in one run:

```
ffmpeg -hide_banner -i fixtures/source.mp4 -i fixtures/excerpt.mp4 \
  -filter_complex "signature=nb_inputs=2:detectmode=full" -f null - 2>&1 | grep -i match
```

Expected shape of the output (offsets will differ):

```
[Parsed_signature_0 @ 0x...] matching of video 0 at 7.000000 and 1 at 0.000000, 150 frames matching
```

It found the excerpt inside the source and told you where. Run the same against `unrelated.mp4` and you get no matching line at all. Run it against `dup_reencode.mp4` and you get `whole video matching`.

Wrap it:

``` python
# dedup/stage3.py
import re, subprocess
from pathlib import Path

MATCH = re.compile(r"matching of video 0 at ([\d.]+) and 1 at ([\d.]+), (\d+) frames matching")
WHOLE = re.compile(r"whole video matching")

def signature_compare(a: Path, b: Path) -> dict | None:
    cmd = ["ffmpeg", "-hide_banner", "-nostats", "-i", str(a), "-i", str(b),
           "-filter_complex", "signature=nb_inputs=2:detectmode=full:th_di=50",
           "-f", "null", "-"]
    out = subprocess.run(cmd, capture_output=True, text=True).stderr
    m = MATCH.search(out)
    if not m:
        return None
    return {
        "offset_a": float(m.group(1)),
        "offset_b": float(m.group(2)),
        "frames": int(m.group(3)),
        "whole": bool(WHOLE.search(out)),
    }
```

`th_di=50` sets the minimum matching sequence to 50 frames (two seconds at 25 fps) so a single similar frame doesn't count. The filter's other thresholds (`th_d`, `th_dc`, `th_xh`, `th_it`) have sane defaults; leave them until you have a labelled set to tune against.

`frames` alone isn't a verdict. You have to divide by something, and the choice is a policy, not a constant.

| Denominator | 6s excerpt in 20s source | Answers | 
|---|---|---|
| source frames (500) | 150 / 500 = 30% | "Is most of my video in theirs?" | 
| shorter video's frames (150) | 150 / 150 = 100% | "Is one of these inside the other?" | 

For dedup you want the second. Most hand-rolled implementations use the first because it's the natural thing to write, and then wonder why excerpts never get flagged.

``` php
# dedup/verdict.py
def coverage(frames: int, frames_a: int, frames_b: int) -> float:
    return frames / min(frames_a, frames_b)
```

Get frame counts from `ffprobe -v error -count_frames -select_streams v:0 -show_entries stream=nb_read_frames -of csv=p=0 file.mp4`, or cheaper, `duration * r_frame_rate` from a plain probe.

``` python
# dedup/worker.py
import sqlite3
from dataclasses import dataclass
from pathlib import Path
from .stage1 import sha256
from .stage2 import phash, hamming
from .stage3 import signature_compare
from .verdict import coverage

PHASH_MAX_DISTANCE = 6
MIN_COVERAGE = 0.4

@dataclass
class Verdict:
    kind: str            # exact | near | contains | new
    match_id: int | None = None
    detail: dict | None = None

def check(path: Path, db: sqlite3.Connection, frames_of) -> Verdict:
    digest = sha256(path)
    row = db.execute("SELECT id FROM assets WHERE sha256=?", (digest,)).fetchone()
    if row:
        return Verdict("exact", row[0])

    h = phash(path)
    candidates = []
    for aid, other_hash, other_path in db.execute("SELECT id, phash, path FROM assets"):
        d = hamming(h, int(other_hash))
        if d <= PHASH_MAX_DISTANCE:
            return Verdict("near", aid, {"hamming": d})
        if d <= 20:                          # loose band: worth a stage-3 look
            candidates.append((aid, Path(other_path)))

    for aid, other in candidates:
        m = signature_compare(other, path)
        if m and coverage(m["frames"], frames_of(other), frames_of(path)) >= MIN_COVERAGE:
            return Verdict("contains", aid, m)

    db.execute("INSERT INTO assets(path, sha256, phash) VALUES (?,?,?)",
               (str(path), digest, str(h)))
    db.commit()
    return Verdict("new")
```

Schema:

```
-- dedup/schema.sql
CREATE TABLE IF NOT EXISTS assets (
  id     INTEGER PRIMARY KEY,
  path   TEXT NOT NULL,
  sha256 TEXT NOT NULL UNIQUE,
  phash  TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_phash ON assets(phash);
```

💡 Tip: the linear scan over `phash` is fine for tens of thousands of assets. Past that, bucket by the top 16 bits of the hash, or move to a BK-tree, before you reach for a vector database.

Run it over the fixtures in order:

```
python -m dedup.cli fixtures/source.mp4 fixtures/dup_reencode.mp4 \
                    fixtures/dup_watermark.mp4 fixtures/excerpt.mp4 fixtures/unrelated.mp4
# source.mp4        new
# dup_reencode.mp4  near      (hamming=<small>)
# dup_watermark.mp4 near      (hamming=<small>)
# excerpt.mp4       contains  (offset_a≈7.0, frames≈150, coverage≈1.00)
# unrelated.mp4     new
# exact distances and frame counts depend on your FFmpeg build; the verdicts should not
```

`.sig` once and keep it next to the media; stage 3 becomes comparison-only.`near` and `contains` rows by hand before you let the worker skip an encode.`check()` into your upload webhook or queue consumer so it runs `reference` table for takedown assets and run stage 3 against it on every upload, not just nightly.`ffmpeg -ss <offset_a> -i source.mp4 -frames:v 1 a.png` next to the same from the candidate is a five-second visual sanity check.
Tag it #python and #ffmpeg if you write up your own thresholds; the tuning data is the part nobody shares.
