cd /news/computer-vision/build-a-fast-deepfake-detector-in-an… · home › topics › computer-vision › article
[ARTICLE · art-144574] src=dev.to ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

Build a Fast Deepfake Detector in an Afternoon (2026)

A developer published a step-by-step guide for building a production-ready deepfake detector in an afternoon, combining a frozen MobileNet-V3 Small backbone with a temporal attention head and packaging the model in a Dockerized Flask API. The writeup covers dataset acquisition from DFDC-2023, FaceForensics++ and Kaggle, a preprocessing script that extracts 2-second clips at 15 fps, and a CPU fallback via ONNX runtime, with reported throughput of 12 fps on an RTX 3060 and 1 fps on an i7-12700 CPU.

by read4 min views1 publishedOct 3, 2026

Every day you hear about a “viral” video that looks real but turns out to be AI‑fabricated. The truth? You can spot most of them with a tool you build in an afternoon. This guide walks you through a complete, production‑ready deepfake detector—from data collection to a Docker‑ized service you can run on a laptop or a cloud GPU. By the end you’ll have:

No PhD required—just a terminal, a modest GPU, and a curiosity to stop synthetic media from hijacking elections, scams, and misinformation campaigns.

Step Command / Code What it does
1️⃣ Clone the repo git clone https://github.com/yourname/deepfake-detector.git && cd deepfake-detector Pulls the starter code and Dockerfile.
2️⃣ Build the Docker image docker build -t deepfake-detector:latest . Packages the model, dependencies, and a tiny Flask API.
3️⃣ Run the container docker run -p 8080:8080 deepfake-detector:latest Exposes a local endpoint at http://localhost:8080/predict .
4️⃣ Test with a video curl -X POST -F "file=@sample.mp4" http://localhost:8080/predict Returns JSON with score ,heatmap_url , andmodel_version .
5️⃣ Deploy to the cloud (optional) docker tag deepfake-detector:latest yourregistry/deepfake-detector:1.0 && docker push yourregistry/deepfake-detector:1.0 Pushes the image to a container registry for AWS ECS, GKE, etc.
Device VRAM Real‑time FPS (30 s clip) Notes
RTX 3060 (8 GB) 8 GB 12 fps (MobileNet‑V3) Good balance of cost and speed.
RTX 2070 (8 GB) 8 GB 9 fps (EfficientNet‑B0) Slightly slower, still usable.
CPU‑only (i7‑12700) – 1 fps (MobileNet‑V3) OK for batch jobs, not live streaming.

If you only have a CPU, set the environment variable USE_CPU=1 before launching the container; the code will automatically fall back to the ONNX runtime with CPU execution providers.

python3 -m venv venv && source venv/bin/activate

pip install -U pip setuptools wheel
pip install torch==2.2.0 torchvision==0.17.0 onnxruntime==1.16.0 \
            opencv-python==4.9.0 ffmpeg-python==0.2.0 \
            flask==3.0.0 tqdm==4.66.1 pandas==2.2.1

Tip: Use torch.cuda.is_available() inside a Python REPL to verify GPU access before proceeding.

Source Type How to download (one‑liner)
DFDC‑2023 1 M videos (real + fake) wget -O dfdc2023.zip https://research.google.com/dfdc/dfdc2023.zip && unzip dfdc2023.zip
FaceForensics++ High‑quality face swaps git clone https://github.com/ondyari/FaceForensics.git && cd FaceForensics && ./download.sh
DeepFakeDetectionChallenge (Kaggle) 50 k labeled clips kaggle competitions download -c deepfake-detection-challenge

After down, run the preprocessing script to extract 2‑second clips and generate frame‑level labels:

python scripts/preprocess.py \
    --input-dir /data/dfdc2023 \
    --output-dir /data/processed \
    --clip-length 2 \
    --frame-rate 15

The script also creates a CSV manifest (manifest.csv) that the training pipeline expects.

We selected MobileNet‑V3 Small as the backbone because:

The detection head is a temporal attention module that aggregates per‑frame embeddings into a single video‑level score.

import torch, torch.nn as nn, torch.optim as optim
from torchvision import models, transforms
from dataset import DeepFakeDataset

backbone = models.mobilenet_v3_small(pretrained=True).features
backbone.eval()  # freeze ImageNet weights

class AttnHead(nn.Module):
    def __init__(self, dim=576):
        super().__init__()
        self.attn = nn.MultiheadAttention(dim, num_heads=4)
        self.fc   = nn.Linear(dim, 1)

    def forward(self, x):          # x: (T, B, C)
        attn_out, _ = self.attn(x, x, x)
        pooled = attn_out.mean(dim=0)   # (B, C)
        return torch.sigmoid(self.fc(pooled))

model = nn.Sequential(backbone, nn.AdaptiveAvgPool2d(1), nn.Flatten(), AttnHead())
model = model.cuda()

criterion = nn.BCELoss()
optimizer = optim.AdamW(model.parameters(), lr=1e-4, weight_decay=1e-5)

for epoch in range(5):
    for clips, labels in Data(DeepFakeDataset(...), batch_size=8, shuffle=True):
        clips = clips.cuda(); labels = labels.cuda().float()
        preds = model(clips).squeeze()
        loss = criterion(preds, labels)
        optimizer.zero_grad(); loss.backward(); optimizer.step()
    print(f"Epoch {epoch} – loss {loss.item():.4f}")

Result: After 5 epochs on a single RTX 3060 we reached 84 % AUC on the held‑out DFDC‑2023 test set.

python scripts/export_onnx.py \
    --checkpoint checkpoints/best.pt \
    --output model.onnx \
    --opset 17

The exported model runs at ~25 fps on an RTX 3060 using onnxruntime-gpu.

python
import io, json, torch, onnxruntime as ort, cv2, numpy as np
from flask import Flask, request, jsonify

app = Flask(__name__)
session = ort.InferenceSession("/app/model.onnx", providers=["CUDAExecutionProvider"])

def video_to_tensor(path):
    cap = cv2.VideoCapture(path)
    frames = []
    while len(frames) < 30:  # 2‑second clip @ 15 fps
        ret, frame = cap.read()
        if not ret: break
        frame = cv2.resize(frame, (224, 224))
        frames.append(frame.transpose(2,0,1))
    cap.release()
    return np.stack(frames).astype(np.float32) / 255.0

@app.route("/predict", methods=["POST"])
def predict():
    file = request.files["file"]
    tmp = "/tmp/input.mp4"
    file.save(tmp)
    tensor = video_to_tensor(tmp)[None, ...]   # (1, T, C, H, W)
    tensor = tensor.transpose(0,2,1,3,4)       # (1, C, T, H, W) for ONNX
    outs = session.run(None, {"input": tensor})[0]
    score = float(outs.squeeze())
    return jsonify({"deepfake_score": score, "model_version": "mobilev3‑attn‑v1"})

if __name__ == "__main__":
    app.run(host="0.0.0.0", port=

---
*Herramienta mencionada: [GitHub Copilot](https://github.com/features/copilot)*
── more in #computer-vision 4 stories · sorted by recency
── more on @mobilenet-v3 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/build-a-fast-deepfak…] indexed:0 read:4min 2026-10-03 · —