cd /news/computer-vision/build-a-three-stage-video-dedup-work… · home › topics › computer-vision › article
[ARTICLE · art-141592] src=dev.to ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

Build a three-stage video dedup worker: SHA-256, perceptual hash, then FFmpeg's MPEG-7 signature

A developer built a three-stage video deduplication worker in Python that combines SHA-256 byte hashing, a 64-bit perceptual hash from the videohash2 library checked via Hamming distance, and FFmpeg's MPEG-7 signature filter for containment detection. The staged design runs the cheap byte and perceptual checks on every upload and reserves the expensive signature comparison for candidates, returning exact, near, contains, or new verdicts backed by a SQLite index. The author notes the perceptual stage cannot detect that one video is an excerpt of another, which is why the MPEG-7 signature stage exists.

by read7 min views2 publishedSep 29, 2026

TL;DR #

We'll build a dedup worker that catches re-encoded, resized, watermarked and trimmed copies of videos you already have. Stage 1 is a byte hash, stage 2 is a 64-bit perceptual hash checked by Hamming distance, stage 3 is FFmpeg's built-in MPEG-7 signature filter for "does A contain B, and where". Python 3.12, FFmpeg 7 or newer, SQLite.

A dedup.py module with one entry point, check(path) -> Verdict, that returns exact, near, contains, or new, plus a SQLite index so it gets faster as the library grows. No ML, no GPUs, nothing that isn't pip install or apt install.

The reason for three stages instead of one: the cheap check should run on every upload before you spend money transcoding, and the expensive check should only run on candidates the cheap check surfaced.

sudo apt-get install -y ffmpeg     # 7.x or 8.x; the signature filter has shipped since 2017
python3 -m venv .venv && source .venv/bin/activate
pip install videohash2            # maintained fork of videohash; pin the current release in requirements.txt
ffmpeg -hide_banner -filters | grep signature

You should see:

 ... signature          N->V       Calculate the MPEG-7 video signature

If that line is missing, your FFmpeg build was configured without it; grab a static build.

We'll synthesise a source and three "duplicates" so the tests are reproducible and the repo stays binary-free.

set -e
mkdir -p fixtures && cd fixtures

ffmpeg -y -f lavfi -i "testsrc2=duration=20:size=640x360:rate=25,format=yuv420p" \
       -f lavfi -i "sine=frequency=440:duration=20" \
       -c:v libx264 -crf 20 -c:a aac -shortest source.mp4

ffmpeg -y -i source.mp4 -vf scale=320:180 -c:v libx264 -crf 32 -c:a aac dup_reencode.mp4

ffmpeg -y -i source.mp4 -vf "drawbox=x=10:y=10:w=80:h=30:color=white@0.8:t=fill" \
       -c:v libx264 -crf 23 -c:a copy dup_watermark.mp4

ffmpeg -y -ss 7 -i source.mp4 -t 6 -c:v libx264 -crf 23 -c:a aac excerpt.mp4

ffmpeg -y -f lavfi -i "mandelbrot=size=640x360:rate=25" -t 20 \
       -c:v libx264 -crf 23 unrelated.mp4

sha256sum fixtures/*.mp4 gives you four different hashes for what a human would call two videos. That's the problem.

import hashlib
from pathlib import Path

def sha256(path: Path, chunk: int = 1 << 20) -> str:
    h = hashlib.sha256()
    with path.open("rb") as f:
        while blk := f.read(chunk):
            h.update(blk)
    return h.hexdigest()

This catches the accidental double upload and nothing else. It stays because it costs nothing and it's the only stage with zero false positives.

videohash2 samples one frame per second, shrinks each to 144x144, tiles them into a collage, and wavelet-hashes the collage into 64 bits. Two re-encodes of the same clip land a few bits apart.

from pathlib import Path
from videohash2 import VideoHash

def phash(path: Path) -> int:
    vh = VideoHash(path=str(path))
    return int(vh.hash_hex, 16)          # 64-bit int

def hamming(a: int, b: int) -> int:
    return (a ^ b).bit_count()

Try it on the fixtures:

>>> from pathlib import Path
>>> from dedup.stage2 import phash, hamming
>>> src = phash(Path("fixtures/source.mp4"))
>>> for name in ["dup_reencode", "dup_watermark", "excerpt", "unrelated"]:
...     print(name, hamming(src, phash(Path(f"fixtures/{name}.mp4"))))

You'll see the re-encode and the watermark land close to zero and unrelated land far away. The interesting one is excerpt: it will not be close. That's by design and the library says so in its docs: a whole-video collage hash cannot tell you that one video is a part of another. That's what stage 3 is for.

⚠️ Note: pick the Hamming threshold conservatively and log near-misses. A false positive here means one user's upload gets served as another user's asset. Start tight (single digits), review what you're rejecting, and loosen only with evidence.

FFmpeg's signature filter fingerprints every frame and, given two or more inputs, searches for matching sequences. First, generate and store a signature per asset. Generation is the slow part; the maintainer of one tool built on this filter reports 19 minutes to fingerprint 96 files versus 45 seconds to compare them all. Never regenerate.

ffmpeg -hide_banner -i fixtures/source.mp4 \
  -vf "signature=format=binary:filename=fixtures/source.sig" -f null -

Then compare two inputs in one run:

ffmpeg -hide_banner -i fixtures/source.mp4 -i fixtures/excerpt.mp4 \
  -filter_complex "signature=nb_inputs=2:detectmode=full" -f null - 2>&1 | grep -i match

Expected shape of the output (offsets will differ):

[Parsed_signature_0 @ 0x...] matching of video 0 at 7.000000 and 1 at 0.000000, 150 frames matching

It found the excerpt inside the source and told you where. Run the same against unrelated.mp4 and you get no matching line at all. Run it against dup_reencode.mp4 and you get whole video matching.

Wrap it:

import re, subprocess
from pathlib import Path

MATCH = re.compile(r"matching of video 0 at ([\d.]+) and 1 at ([\d.]+), (\d+) frames matching")
WHOLE = re.compile(r"whole video matching")

def signature_compare(a: Path, b: Path) -> dict | None:
    cmd = ["ffmpeg", "-hide_banner", "-nostats", "-i", str(a), "-i", str(b),
           "-filter_complex", "signature=nb_inputs=2:detectmode=full:th_di=50",
           "-f", "null", "-"]
    out = subprocess.run(cmd, capture_output=True, text=True).stderr
    m = MATCH.search(out)
    if not m:
        return None
    return {
        "offset_a": float(m.group(1)),
        "offset_b": float(m.group(2)),
        "frames": int(m.group(3)),
        "whole": bool(WHOLE.search(out)),
    }

th_di=50 sets the minimum matching sequence to 50 frames (two seconds at 25 fps) so a single similar frame doesn't count. The filter's other thresholds (th_d, th_dc, th_xh, th_it) have sane defaults; leave them until you have a labelled set to tune against.

frames alone isn't a verdict. You have to divide by something, and the choice is a policy, not a constant.

Denominator 6s excerpt in 20s source Answers
source frames (500) 150 / 500 = 30% "Is most of my video in theirs?"
shorter video's frames (150) 150 / 150 = 100% "Is one of these inside the other?"

For dedup you want the second. Most hand-rolled implementations use the first because it's the natural thing to write, and then wonder why excerpts never get flagged.

def coverage(frames: int, frames_a: int, frames_b: int) -> float:
    return frames / min(frames_a, frames_b)

Get frame counts from ffprobe -v error -count_frames -select_streams v:0 -show_entries stream=nb_read_frames -of csv=p=0 file.mp4, or cheaper, duration * r_frame_rate from a plain probe.

import sqlite3
from dataclasses import dataclass
from pathlib import Path
from .stage1 import sha256
from .stage2 import phash, hamming
from .stage3 import signature_compare
from .verdict import coverage

PHASH_MAX_DISTANCE = 6
MIN_COVERAGE = 0.4

@dataclass
class Verdict:
    kind: str            # exact | near | contains | new
    match_id: int | None = None
    detail: dict | None = None

def check(path: Path, db: sqlite3.Connection, frames_of) -> Verdict:
    digest = sha256(path)
    row = db.execute("SELECT id FROM assets WHERE sha256=?", (digest,)).fetchone()
    if row:
        return Verdict("exact", row[0])

    h = phash(path)
    candidates = []
    for aid, other_hash, other_path in db.execute("SELECT id, phash, path FROM assets"):
        d = hamming(h, int(other_hash))
        if d <= PHASH_MAX_DISTANCE:
            return Verdict("near", aid, {"hamming": d})
        if d <= 20:                          # loose band: worth a stage-3 look
            candidates.append((aid, Path(other_path)))

    for aid, other in candidates:
        m = signature_compare(other, path)
        if m and coverage(m["frames"], frames_of(other), frames_of(path)) >= MIN_COVERAGE:
            return Verdict("contains", aid, m)

    db.execute("INSERT INTO assets(path, sha256, phash) VALUES (?,?,?)",
               (str(path), digest, str(h)))
    db.commit()
    return Verdict("new")

Schema:

-- dedup/schema.sql
CREATE TABLE IF NOT EXISTS assets (
  id     INTEGER PRIMARY KEY,
  path   TEXT NOT NULL,
  sha256 TEXT NOT NULL UNIQUE,
  phash  TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_phash ON assets(phash);

💡 Tip: the linear scan over phash is fine for tens of thousands of assets. Past that, bucket by the top 16 bits of the hash, or move to a BK-tree, before you reach for a vector database.

Run it over the fixtures in order:

python -m dedup.cli fixtures/source.mp4 fixtures/dup_reencode.mp4 \
                    fixtures/dup_watermark.mp4 fixtures/excerpt.mp4 fixtures/unrelated.mp4

.sig once and keep it next to the media; stage 3 becomes comparison-only.near and contains rows by hand before you let the worker skip an encode.check() into your upload webhook or queue consumer so it runs reference table for takedown assets and run stage 3 against it on every upload, not just nightly.ffmpeg -ss <offset_a> -i source.mp4 -frames:v 1 a.png next to the same from the candidate is a five-second visual sanity check. Tag it #python and #ffmpeg if you write up your own thresholds; the tuning data is the part nobody shares.

── more in #computer-vision 4 stories · sorted by recency
── more on @ffmpeg 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/build-a-three-stage-…] indexed:0 read:7min 2026-09-29 · —