{"slug": "build-a-three-stage-video-dedup-worker-sha-256-perceptual-hash-then-ffmpeg-s-7", "title": "Build a three-stage video dedup worker: SHA-256, perceptual hash, then FFmpeg's MPEG-7 signature", "summary": "A developer built a three-stage video deduplication worker in Python that combines SHA-256 byte hashing, a 64-bit perceptual hash from the videohash2 library checked via Hamming distance, and FFmpeg's MPEG-7 signature filter for containment detection. The staged design runs the cheap byte and perceptual checks on every upload and reserves the expensive signature comparison for candidates, returning exact, near, contains, or new verdicts backed by a SQLite index. The author notes the perceptual stage cannot detect that one video is an excerpt of another, which is why the MPEG-7 signature stage exists.", "body_md": "## TL;DR\n\nWe'll build a dedup worker that catches re-encoded, resized, watermarked and trimmed copies of videos you already have. Stage 1 is a byte hash, stage 2 is a 64-bit perceptual hash checked by Hamming distance, stage 3 is FFmpeg's built-in MPEG-7 `signature` filter for \"does A contain B, and where\". Python 3.12, FFmpeg 7 or newer, SQLite.\n\nA `dedup.py` module with one entry point, `check(path) -> Verdict`, that returns `exact`, `near`, `contains`, or `new`, plus a SQLite index so it gets faster as the library grows. No ML, no GPUs, nothing that isn't `pip install` or `apt install`.\n\nThe reason for three stages instead of one: the cheap check should run on every upload before you spend money transcoding, and the expensive check should only run on candidates the cheap check surfaced.\n\n```\nsudo apt-get install -y ffmpeg     # 7.x or 8.x; the signature filter has shipped since 2017\npython3 -m venv .venv && source .venv/bin/activate\npip install videohash2            # maintained fork of videohash; pin the current release in requirements.txt\nffmpeg -hide_banner -filters | grep signature\n```\n\nYou should see:\n\n``` php\n ... signature          N->V       Calculate the MPEG-7 video signature\n```\n\nIf that line is missing, your FFmpeg build was configured without it; grab a static build.\n\nWe'll synthesise a source and three \"duplicates\" so the tests are reproducible and the repo stays binary-free.\n\n```\n# scripts/make_fixtures.sh\nset -e\nmkdir -p fixtures && cd fixtures\n\n# 20-second source\nffmpeg -y -f lavfi -i \"testsrc2=duration=20:size=640x360:rate=25,format=yuv420p\" \\\n       -f lavfi -i \"sine=frequency=440:duration=20\" \\\n       -c:v libx264 -crf 20 -c:a aac -shortest source.mp4\n\n# A: re-encode at lower quality and resolution\nffmpeg -y -i source.mp4 -vf scale=320:180 -c:v libx264 -crf 32 -c:a aac dup_reencode.mp4\n\n# B: watermark in the corner\nffmpeg -y -i source.mp4 -vf \"drawbox=x=10:y=10:w=80:h=30:color=white@0.8:t=fill\" \\\n       -c:v libx264 -crf 23 -c:a copy dup_watermark.mp4\n\n# C: 6-second excerpt starting at 0:07 (containment case)\nffmpeg -y -ss 7 -i source.mp4 -t 6 -c:v libx264 -crf 23 -c:a aac excerpt.mp4\n\n# D: unrelated video\nffmpeg -y -f lavfi -i \"mandelbrot=size=640x360:rate=25\" -t 20 \\\n       -c:v libx264 -crf 23 unrelated.mp4\n```\n\n`sha256sum fixtures/*.mp4` gives you four different hashes for what a human would call two videos. That's the problem.\n\n``` python\n# dedup/stage1.py\nimport hashlib\nfrom pathlib import Path\n\ndef sha256(path: Path, chunk: int = 1 << 20) -> str:\n    h = hashlib.sha256()\n    with path.open(\"rb\") as f:\n        while blk := f.read(chunk):\n            h.update(blk)\n    return h.hexdigest()\n```\n\nThis catches the accidental double upload and nothing else. It stays because it costs nothing and it's the only stage with zero false positives.\n\n`videohash2` samples one frame per second, shrinks each to 144x144, tiles them into a collage, and wavelet-hashes the collage into 64 bits. Two re-encodes of the same clip land a few bits apart.\n\n``` python\n# dedup/stage2.py\nfrom pathlib import Path\nfrom videohash2 import VideoHash\n\ndef phash(path: Path) -> int:\n    vh = VideoHash(path=str(path))\n    return int(vh.hash_hex, 16)          # 64-bit int\n\ndef hamming(a: int, b: int) -> int:\n    return (a ^ b).bit_count()\n```\n\nTry it on the fixtures:\n\n``` python\n>>> from pathlib import Path\n>>> from dedup.stage2 import phash, hamming\n>>> src = phash(Path(\"fixtures/source.mp4\"))\n>>> for name in [\"dup_reencode\", \"dup_watermark\", \"excerpt\", \"unrelated\"]:\n...     print(name, hamming(src, phash(Path(f\"fixtures/{name}.mp4\"))))\n```\n\nYou'll see the re-encode and the watermark land close to zero and `unrelated` land far away. The interesting one is `excerpt`: it will **not** be close. That's by design and the library says so in its docs: a whole-video collage hash cannot tell you that one video is a part of another. That's what stage 3 is for.\n\n⚠️ Note: pick the Hamming threshold conservatively and log near-misses. A false positive here means one user's upload gets served as another user's asset. Start tight (single digits), review what you're rejecting, and loosen only with evidence.\n\nFFmpeg's `signature` filter fingerprints every frame and, given two or more inputs, searches for matching sequences. First, generate and **store** a signature per asset. Generation is the slow part; the maintainer of one tool built on this filter reports 19 minutes to fingerprint 96 files versus 45 seconds to compare them all. Never regenerate.\n\n```\n# one-time per asset: write a binary signature next to it\nffmpeg -hide_banner -i fixtures/source.mp4 \\\n  -vf \"signature=format=binary:filename=fixtures/source.sig\" -f null -\n```\n\nThen compare two inputs in one run:\n\n```\nffmpeg -hide_banner -i fixtures/source.mp4 -i fixtures/excerpt.mp4 \\\n  -filter_complex \"signature=nb_inputs=2:detectmode=full\" -f null - 2>&1 | grep -i match\n```\n\nExpected shape of the output (offsets will differ):\n\n```\n[Parsed_signature_0 @ 0x...] matching of video 0 at 7.000000 and 1 at 0.000000, 150 frames matching\n```\n\nIt found the excerpt inside the source and told you where. Run the same against `unrelated.mp4` and you get no matching line at all. Run it against `dup_reencode.mp4` and you get `whole video matching`.\n\nWrap it:\n\n``` python\n# dedup/stage3.py\nimport re, subprocess\nfrom pathlib import Path\n\nMATCH = re.compile(r\"matching of video 0 at ([\\d.]+) and 1 at ([\\d.]+), (\\d+) frames matching\")\nWHOLE = re.compile(r\"whole video matching\")\n\ndef signature_compare(a: Path, b: Path) -> dict | None:\n    cmd = [\"ffmpeg\", \"-hide_banner\", \"-nostats\", \"-i\", str(a), \"-i\", str(b),\n           \"-filter_complex\", \"signature=nb_inputs=2:detectmode=full:th_di=50\",\n           \"-f\", \"null\", \"-\"]\n    out = subprocess.run(cmd, capture_output=True, text=True).stderr\n    m = MATCH.search(out)\n    if not m:\n        return None\n    return {\n        \"offset_a\": float(m.group(1)),\n        \"offset_b\": float(m.group(2)),\n        \"frames\": int(m.group(3)),\n        \"whole\": bool(WHOLE.search(out)),\n    }\n```\n\n`th_di=50` sets the minimum matching sequence to 50 frames (two seconds at 25 fps) so a single similar frame doesn't count. The filter's other thresholds (`th_d`, `th_dc`, `th_xh`, `th_it`) have sane defaults; leave them until you have a labelled set to tune against.\n\n`frames` alone isn't a verdict. You have to divide by something, and the choice is a policy, not a constant.\n\n| Denominator | 6s excerpt in 20s source | Answers | \n|---|---|---|\n| source frames (500) | 150 / 500 = 30% | \"Is most of my video in theirs?\" | \n| shorter video's frames (150) | 150 / 150 = 100% | \"Is one of these inside the other?\" | \n\nFor dedup you want the second. Most hand-rolled implementations use the first because it's the natural thing to write, and then wonder why excerpts never get flagged.\n\n``` php\n# dedup/verdict.py\ndef coverage(frames: int, frames_a: int, frames_b: int) -> float:\n    return frames / min(frames_a, frames_b)\n```\n\nGet frame counts from `ffprobe -v error -count_frames -select_streams v:0 -show_entries stream=nb_read_frames -of csv=p=0 file.mp4`, or cheaper, `duration * r_frame_rate` from a plain probe.\n\n``` python\n# dedup/worker.py\nimport sqlite3\nfrom dataclasses import dataclass\nfrom pathlib import Path\nfrom .stage1 import sha256\nfrom .stage2 import phash, hamming\nfrom .stage3 import signature_compare\nfrom .verdict import coverage\n\nPHASH_MAX_DISTANCE = 6\nMIN_COVERAGE = 0.4\n\n@dataclass\nclass Verdict:\n    kind: str            # exact | near | contains | new\n    match_id: int | None = None\n    detail: dict | None = None\n\ndef check(path: Path, db: sqlite3.Connection, frames_of) -> Verdict:\n    digest = sha256(path)\n    row = db.execute(\"SELECT id FROM assets WHERE sha256=?\", (digest,)).fetchone()\n    if row:\n        return Verdict(\"exact\", row[0])\n\n    h = phash(path)\n    candidates = []\n    for aid, other_hash, other_path in db.execute(\"SELECT id, phash, path FROM assets\"):\n        d = hamming(h, int(other_hash))\n        if d <= PHASH_MAX_DISTANCE:\n            return Verdict(\"near\", aid, {\"hamming\": d})\n        if d <= 20:                          # loose band: worth a stage-3 look\n            candidates.append((aid, Path(other_path)))\n\n    for aid, other in candidates:\n        m = signature_compare(other, path)\n        if m and coverage(m[\"frames\"], frames_of(other), frames_of(path)) >= MIN_COVERAGE:\n            return Verdict(\"contains\", aid, m)\n\n    db.execute(\"INSERT INTO assets(path, sha256, phash) VALUES (?,?,?)\",\n               (str(path), digest, str(h)))\n    db.commit()\n    return Verdict(\"new\")\n```\n\nSchema:\n\n```\n-- dedup/schema.sql\nCREATE TABLE IF NOT EXISTS assets (\n  id     INTEGER PRIMARY KEY,\n  path   TEXT NOT NULL,\n  sha256 TEXT NOT NULL UNIQUE,\n  phash  TEXT NOT NULL\n);\nCREATE INDEX IF NOT EXISTS idx_phash ON assets(phash);\n```\n\n💡 Tip: the linear scan over `phash` is fine for tens of thousands of assets. Past that, bucket by the top 16 bits of the hash, or move to a BK-tree, before you reach for a vector database.\n\nRun it over the fixtures in order:\n\n```\npython -m dedup.cli fixtures/source.mp4 fixtures/dup_reencode.mp4 \\\n                    fixtures/dup_watermark.mp4 fixtures/excerpt.mp4 fixtures/unrelated.mp4\n# source.mp4        new\n# dup_reencode.mp4  near      (hamming=<small>)\n# dup_watermark.mp4 near      (hamming=<small>)\n# excerpt.mp4       contains  (offset_a≈7.0, frames≈150, coverage≈1.00)\n# unrelated.mp4     new\n# exact distances and frame counts depend on your FFmpeg build; the verdicts should not\n```\n\n`.sig` once and keep it next to the media; stage 3 becomes comparison-only.`near` and `contains` rows by hand before you let the worker skip an encode.`check()` into your upload webhook or queue consumer so it runs `reference` table for takedown assets and run stage 3 against it on every upload, not just nightly.`ffmpeg -ss <offset_a> -i source.mp4 -frames:v 1 a.png` next to the same from the candidate is a five-second visual sanity check.\nTag it #python and #ffmpeg if you write up your own thresholds; the tuning data is the part nobody shares.", "url": "https://wpnews.pro/news/build-a-three-stage-video-dedup-worker-sha-256-perceptual-hash-then-ffmpeg-s-7", "canonical_source": "https://dev.to/masonwritescode/build-a-three-stage-video-dedup-worker-sha-256-perceptual-hash-then-ffmpegs-mpeg-7-signature-578l", "published_at": "2026-09-29 09:12:35+00:00", "updated_at": "2026-09-29 09:16:36.297819+00:00", "lang": "en", "topics": ["computer-vision", "developer-tools", "ai-tools"], "entities": ["FFmpeg", "videohash2", "SQLite", "Python"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/build-a-three-stage-video-dedup-worker-sha-256-perceptual-hash-then-ffmpeg-s-7", "markdown": "https://wpnews.pro/news/build-a-three-stage-video-dedup-worker-sha-256-perceptual-hash-then-ffmpeg-s-7.md", "text": "https://wpnews.pro/news/build-a-three-stage-video-dedup-worker-sha-256-perceptual-hash-then-ffmpeg-s-7.txt", "jsonld": "https://wpnews.pro/news/build-a-three-stage-video-dedup-worker-sha-256-perceptual-hash-then-ffmpeg-s-7.jsonld"}}