Build a three-stage video dedup worker: SHA-256, perceptual hash, then FFmpeg's MPEG-7 signature A developer built a three-stage video deduplication worker in Python that combines SHA-256 byte hashing, a 64-bit perceptual hash from the videohash2 library checked via Hamming distance, and FFmpeg's MPEG-7 signature filter for containment detection. The staged design runs the cheap byte and perceptual checks on every upload and reserves the expensive signature comparison for candidates, returning exact, near, contains, or new verdicts backed by a SQLite index. The author notes the perceptual stage cannot detect that one video is an excerpt of another, which is why the MPEG-7 signature stage exists. TL;DR We'll build a dedup worker that catches re-encoded, resized, watermarked and trimmed copies of videos you already have. Stage 1 is a byte hash, stage 2 is a 64-bit perceptual hash checked by Hamming distance, stage 3 is FFmpeg's built-in MPEG-7 signature filter for "does A contain B, and where". Python 3.12, FFmpeg 7 or newer, SQLite. A dedup.py module with one entry point, check path - Verdict , that returns exact , near , contains , or new , plus a SQLite index so it gets faster as the library grows. No ML, no GPUs, nothing that isn't pip install or apt install . The reason for three stages instead of one: the cheap check should run on every upload before you spend money transcoding, and the expensive check should only run on candidates the cheap check surfaced. sudo apt-get install -y ffmpeg 7.x or 8.x; the signature filter has shipped since 2017 python3 -m venv .venv && source .venv/bin/activate pip install videohash2 maintained fork of videohash; pin the current release in requirements.txt ffmpeg -hide banner -filters | grep signature You should see: php ... signature N- V Calculate the MPEG-7 video signature If that line is missing, your FFmpeg build was configured without it; grab a static build. We'll synthesise a source and three "duplicates" so the tests are reproducible and the repo stays binary-free. scripts/make fixtures.sh set -e mkdir -p fixtures && cd fixtures 20-second source ffmpeg -y -f lavfi -i "testsrc2=duration=20:size=640x360:rate=25,format=yuv420p" \ -f lavfi -i "sine=frequency=440:duration=20" \ -c:v libx264 -crf 20 -c:a aac -shortest source.mp4 A: re-encode at lower quality and resolution ffmpeg -y -i source.mp4 -vf scale=320:180 -c:v libx264 -crf 32 -c:a aac dup reencode.mp4 B: watermark in the corner ffmpeg -y -i source.mp4 -vf "drawbox=x=10:y=10:w=80:h=30:color=white@0.8:t=fill" \ -c:v libx264 -crf 23 -c:a copy dup watermark.mp4 C: 6-second excerpt starting at 0:07 containment case ffmpeg -y -ss 7 -i source.mp4 -t 6 -c:v libx264 -crf 23 -c:a aac excerpt.mp4 D: unrelated video ffmpeg -y -f lavfi -i "mandelbrot=size=640x360:rate=25" -t 20 \ -c:v libx264 -crf 23 unrelated.mp4 sha256sum fixtures/ .mp4 gives you four different hashes for what a human would call two videos. That's the problem. python dedup/stage1.py import hashlib from pathlib import Path def sha256 path: Path, chunk: int = 1 << 20 - str: h = hashlib.sha256 with path.open "rb" as f: while blk := f.read chunk : h.update blk return h.hexdigest This catches the accidental double upload and nothing else. It stays because it costs nothing and it's the only stage with zero false positives. videohash2 samples one frame per second, shrinks each to 144x144, tiles them into a collage, and wavelet-hashes the collage into 64 bits. Two re-encodes of the same clip land a few bits apart. python dedup/stage2.py from pathlib import Path from videohash2 import VideoHash def phash path: Path - int: vh = VideoHash path=str path return int vh.hash hex, 16 64-bit int def hamming a: int, b: int - int: return a ^ b .bit count Try it on the fixtures: python from pathlib import Path from dedup.stage2 import phash, hamming src = phash Path "fixtures/source.mp4" for name in "dup reencode", "dup watermark", "excerpt", "unrelated" : ... print name, hamming src, phash Path f"fixtures/{name}.mp4" You'll see the re-encode and the watermark land close to zero and unrelated land far away. The interesting one is excerpt : it will not be close. That's by design and the library says so in its docs: a whole-video collage hash cannot tell you that one video is a part of another. That's what stage 3 is for. ⚠️ Note: pick the Hamming threshold conservatively and log near-misses. A false positive here means one user's upload gets served as another user's asset. Start tight single digits , review what you're rejecting, and loosen only with evidence. FFmpeg's signature filter fingerprints every frame and, given two or more inputs, searches for matching sequences. First, generate and store a signature per asset. Generation is the slow part; the maintainer of one tool built on this filter reports 19 minutes to fingerprint 96 files versus 45 seconds to compare them all. Never regenerate. one-time per asset: write a binary signature next to it ffmpeg -hide banner -i fixtures/source.mp4 \ -vf "signature=format=binary:filename=fixtures/source.sig" -f null - Then compare two inputs in one run: ffmpeg -hide banner -i fixtures/source.mp4 -i fixtures/excerpt.mp4 \ -filter complex "signature=nb inputs=2:detectmode=full" -f null - 2 &1 | grep -i match Expected shape of the output offsets will differ : Parsed signature 0 @ 0x... matching of video 0 at 7.000000 and 1 at 0.000000, 150 frames matching It found the excerpt inside the source and told you where. Run the same against unrelated.mp4 and you get no matching line at all. Run it against dup reencode.mp4 and you get whole video matching . Wrap it: python dedup/stage3.py import re, subprocess from pathlib import Path MATCH = re.compile r"matching of video 0 at \d. + and 1 at \d. + , \d+ frames matching" WHOLE = re.compile r"whole video matching" def signature compare a: Path, b: Path - dict | None: cmd = "ffmpeg", "-hide banner", "-nostats", "-i", str a , "-i", str b , "-filter complex", "signature=nb inputs=2:detectmode=full:th di=50", "-f", "null", "-" out = subprocess.run cmd, capture output=True, text=True .stderr m = MATCH.search out if not m: return None return { "offset a": float m.group 1 , "offset b": float m.group 2 , "frames": int m.group 3 , "whole": bool WHOLE.search out , } th di=50 sets the minimum matching sequence to 50 frames two seconds at 25 fps so a single similar frame doesn't count. The filter's other thresholds th d , th dc , th xh , th it have sane defaults; leave them until you have a labelled set to tune against. frames alone isn't a verdict. You have to divide by something, and the choice is a policy, not a constant. | Denominator | 6s excerpt in 20s source | Answers | |---|---|---| | source frames 500 | 150 / 500 = 30% | "Is most of my video in theirs?" | | shorter video's frames 150 | 150 / 150 = 100% | "Is one of these inside the other?" | For dedup you want the second. Most hand-rolled implementations use the first because it's the natural thing to write, and then wonder why excerpts never get flagged. php dedup/verdict.py def coverage frames: int, frames a: int, frames b: int - float: return frames / min frames a, frames b Get frame counts from ffprobe -v error -count frames -select streams v:0 -show entries stream=nb read frames -of csv=p=0 file.mp4 , or cheaper, duration r frame rate from a plain probe. python dedup/worker.py import sqlite3 from dataclasses import dataclass from pathlib import Path from .stage1 import sha256 from .stage2 import phash, hamming from .stage3 import signature compare from .verdict import coverage PHASH MAX DISTANCE = 6 MIN COVERAGE = 0.4 @dataclass class Verdict: kind: str exact | near | contains | new match id: int | None = None detail: dict | None = None def check path: Path, db: sqlite3.Connection, frames of - Verdict: digest = sha256 path row = db.execute "SELECT id FROM assets WHERE sha256=?", digest, .fetchone if row: return Verdict "exact", row 0 h = phash path candidates = for aid, other hash, other path in db.execute "SELECT id, phash, path FROM assets" : d = hamming h, int other hash if d <= PHASH MAX DISTANCE: return Verdict "near", aid, {"hamming": d} if d <= 20: loose band: worth a stage-3 look candidates.append aid, Path other path for aid, other in candidates: m = signature compare other, path if m and coverage m "frames" , frames of other , frames of path = MIN COVERAGE: return Verdict "contains", aid, m db.execute "INSERT INTO assets path, sha256, phash VALUES ?,?,? ", str path , digest, str h db.commit return Verdict "new" Schema: -- dedup/schema.sql CREATE TABLE IF NOT EXISTS assets id INTEGER PRIMARY KEY, path TEXT NOT NULL, sha256 TEXT NOT NULL UNIQUE, phash TEXT NOT NULL ; CREATE INDEX IF NOT EXISTS idx phash ON assets phash ; 💡 Tip: the linear scan over phash is fine for tens of thousands of assets. Past that, bucket by the top 16 bits of the hash, or move to a BK-tree, before you reach for a vector database. Run it over the fixtures in order: python -m dedup.cli fixtures/source.mp4 fixtures/dup reencode.mp4 \ fixtures/dup watermark.mp4 fixtures/excerpt.mp4 fixtures/unrelated.mp4 source.mp4 new dup reencode.mp4 near hamming=