TL;DR #
We'll build a dedup worker that catches re-encoded, resized, watermarked and trimmed copies of videos you already have. Stage 1 is a byte hash, stage 2 is a 64-bit perceptual hash checked by Hamming distance, stage 3 is FFmpeg's built-in MPEG-7 signature filter for "does A contain B, and where". Python 3.12, FFmpeg 7 or newer, SQLite.
A dedup.py module with one entry point, check(path) -> Verdict, that returns exact, near, contains, or new, plus a SQLite index so it gets faster as the library grows. No ML, no GPUs, nothing that isn't pip install or apt install.
The reason for three stages instead of one: the cheap check should run on every upload before you spend money transcoding, and the expensive check should only run on candidates the cheap check surfaced.
sudo apt-get install -y ffmpeg # 7.x or 8.x; the signature filter has shipped since 2017
python3 -m venv .venv && source .venv/bin/activate
pip install videohash2 # maintained fork of videohash; pin the current release in requirements.txt
ffmpeg -hide_banner -filters | grep signature
You should see:
... signature N->V Calculate the MPEG-7 video signature
If that line is missing, your FFmpeg build was configured without it; grab a static build.
We'll synthesise a source and three "duplicates" so the tests are reproducible and the repo stays binary-free.
set -e
mkdir -p fixtures && cd fixtures
ffmpeg -y -f lavfi -i "testsrc2=duration=20:size=640x360:rate=25,format=yuv420p" \
-f lavfi -i "sine=frequency=440:duration=20" \
-c:v libx264 -crf 20 -c:a aac -shortest source.mp4
ffmpeg -y -i source.mp4 -vf scale=320:180 -c:v libx264 -crf 32 -c:a aac dup_reencode.mp4
ffmpeg -y -i source.mp4 -vf "drawbox=x=10:y=10:w=80:h=30:color=white@0.8:t=fill" \
-c:v libx264 -crf 23 -c:a copy dup_watermark.mp4
ffmpeg -y -ss 7 -i source.mp4 -t 6 -c:v libx264 -crf 23 -c:a aac excerpt.mp4
ffmpeg -y -f lavfi -i "mandelbrot=size=640x360:rate=25" -t 20 \
-c:v libx264 -crf 23 unrelated.mp4
sha256sum fixtures/*.mp4 gives you four different hashes for what a human would call two videos. That's the problem.
import hashlib
from pathlib import Path
def sha256(path: Path, chunk: int = 1 << 20) -> str:
h = hashlib.sha256()
with path.open("rb") as f:
while blk := f.read(chunk):
h.update(blk)
return h.hexdigest()
This catches the accidental double upload and nothing else. It stays because it costs nothing and it's the only stage with zero false positives.
videohash2 samples one frame per second, shrinks each to 144x144, tiles them into a collage, and wavelet-hashes the collage into 64 bits. Two re-encodes of the same clip land a few bits apart.
from pathlib import Path
from videohash2 import VideoHash
def phash(path: Path) -> int:
vh = VideoHash(path=str(path))
return int(vh.hash_hex, 16) # 64-bit int
def hamming(a: int, b: int) -> int:
return (a ^ b).bit_count()
Try it on the fixtures:
>>> from pathlib import Path
>>> from dedup.stage2 import phash, hamming
>>> src = phash(Path("fixtures/source.mp4"))
>>> for name in ["dup_reencode", "dup_watermark", "excerpt", "unrelated"]:
... print(name, hamming(src, phash(Path(f"fixtures/{name}.mp4"))))
You'll see the re-encode and the watermark land close to zero and unrelated land far away. The interesting one is excerpt: it will not be close. That's by design and the library says so in its docs: a whole-video collage hash cannot tell you that one video is a part of another. That's what stage 3 is for.
⚠️ Note: pick the Hamming threshold conservatively and log near-misses. A false positive here means one user's upload gets served as another user's asset. Start tight (single digits), review what you're rejecting, and loosen only with evidence.
FFmpeg's signature filter fingerprints every frame and, given two or more inputs, searches for matching sequences. First, generate and store a signature per asset. Generation is the slow part; the maintainer of one tool built on this filter reports 19 minutes to fingerprint 96 files versus 45 seconds to compare them all. Never regenerate.
ffmpeg -hide_banner -i fixtures/source.mp4 \
-vf "signature=format=binary:filename=fixtures/source.sig" -f null -
Then compare two inputs in one run:
ffmpeg -hide_banner -i fixtures/source.mp4 -i fixtures/excerpt.mp4 \
-filter_complex "signature=nb_inputs=2:detectmode=full" -f null - 2>&1 | grep -i match
Expected shape of the output (offsets will differ):
[Parsed_signature_0 @ 0x...] matching of video 0 at 7.000000 and 1 at 0.000000, 150 frames matching
It found the excerpt inside the source and told you where. Run the same against unrelated.mp4 and you get no matching line at all. Run it against dup_reencode.mp4 and you get whole video matching.
Wrap it:
import re, subprocess
from pathlib import Path
MATCH = re.compile(r"matching of video 0 at ([\d.]+) and 1 at ([\d.]+), (\d+) frames matching")
WHOLE = re.compile(r"whole video matching")
def signature_compare(a: Path, b: Path) -> dict | None:
cmd = ["ffmpeg", "-hide_banner", "-nostats", "-i", str(a), "-i", str(b),
"-filter_complex", "signature=nb_inputs=2:detectmode=full:th_di=50",
"-f", "null", "-"]
out = subprocess.run(cmd, capture_output=True, text=True).stderr
m = MATCH.search(out)
if not m:
return None
return {
"offset_a": float(m.group(1)),
"offset_b": float(m.group(2)),
"frames": int(m.group(3)),
"whole": bool(WHOLE.search(out)),
}
th_di=50 sets the minimum matching sequence to 50 frames (two seconds at 25 fps) so a single similar frame doesn't count. The filter's other thresholds (th_d, th_dc, th_xh, th_it) have sane defaults; leave them until you have a labelled set to tune against.
frames alone isn't a verdict. You have to divide by something, and the choice is a policy, not a constant.
| Denominator | 6s excerpt in 20s source | Answers |
|---|---|---|
| source frames (500) | 150 / 500 = 30% | "Is most of my video in theirs?" |
| shorter video's frames (150) | 150 / 150 = 100% | "Is one of these inside the other?" |
For dedup you want the second. Most hand-rolled implementations use the first because it's the natural thing to write, and then wonder why excerpts never get flagged.
def coverage(frames: int, frames_a: int, frames_b: int) -> float:
return frames / min(frames_a, frames_b)
Get frame counts from ffprobe -v error -count_frames -select_streams v:0 -show_entries stream=nb_read_frames -of csv=p=0 file.mp4, or cheaper, duration * r_frame_rate from a plain probe.
import sqlite3
from dataclasses import dataclass
from pathlib import Path
from .stage1 import sha256
from .stage2 import phash, hamming
from .stage3 import signature_compare
from .verdict import coverage
PHASH_MAX_DISTANCE = 6
MIN_COVERAGE = 0.4
@dataclass
class Verdict:
kind: str # exact | near | contains | new
match_id: int | None = None
detail: dict | None = None
def check(path: Path, db: sqlite3.Connection, frames_of) -> Verdict:
digest = sha256(path)
row = db.execute("SELECT id FROM assets WHERE sha256=?", (digest,)).fetchone()
if row:
return Verdict("exact", row[0])
h = phash(path)
candidates = []
for aid, other_hash, other_path in db.execute("SELECT id, phash, path FROM assets"):
d = hamming(h, int(other_hash))
if d <= PHASH_MAX_DISTANCE:
return Verdict("near", aid, {"hamming": d})
if d <= 20: # loose band: worth a stage-3 look
candidates.append((aid, Path(other_path)))
for aid, other in candidates:
m = signature_compare(other, path)
if m and coverage(m["frames"], frames_of(other), frames_of(path)) >= MIN_COVERAGE:
return Verdict("contains", aid, m)
db.execute("INSERT INTO assets(path, sha256, phash) VALUES (?,?,?)",
(str(path), digest, str(h)))
db.commit()
return Verdict("new")
Schema:
-- dedup/schema.sql
CREATE TABLE IF NOT EXISTS assets (
id INTEGER PRIMARY KEY,
path TEXT NOT NULL,
sha256 TEXT NOT NULL UNIQUE,
phash TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_phash ON assets(phash);
💡 Tip: the linear scan over phash is fine for tens of thousands of assets. Past that, bucket by the top 16 bits of the hash, or move to a BK-tree, before you reach for a vector database.
Run it over the fixtures in order:
python -m dedup.cli fixtures/source.mp4 fixtures/dup_reencode.mp4 \
fixtures/dup_watermark.mp4 fixtures/excerpt.mp4 fixtures/unrelated.mp4
.sig once and keep it next to the media; stage 3 becomes comparison-only.near and contains rows by hand before you let the worker skip an encode.check() into your upload webhook or queue consumer so it runs reference table for takedown assets and run stage 3 against it on every upload, not just nightly.ffmpeg -ss <offset_a> -i source.mp4 -frames:v 1 a.png next to the same from the candidate is a five-second visual sanity check.
Tag it #python and #ffmpeg if you write up your own thresholds; the tuning data is the part nobody shares.