Every few months someone ships a new way to detect AI-generated text or images, and every few months someone else shows how to strip it. Paraphrase the text, crop or re-encode the image, and the hidden signal gets weaker or disappears. Detection is a cat and mouse game, and the mouse only needs to win once.
There is a calmer way to think about this. Instead of asking "can I prove this was made by AI?", ask "can I prove where this came from and that nobody changed it since?" That is provenance, and it is a much easier problem to engineer.
| Watermark / detector | Signed provenance | |
|---|---|---|
| Where it lives | hidden inside the content | a small record next to the content |
| What it claims | "this probably came from model X" | "this exact file came from source Y at time T" |
| Survives edits? | degrades with paraphrase, crops, re-encoding | any edit breaks the signature, which is the point |
| Failure mode | silent false negatives and false positives | missing or invalid signature, clearly visible |
| Who can check | usually only the vendor | anyone with the public key |
A watermark tries to make the content itself carry the evidence. Provenance keeps the evidence outside the content and makes it cryptographically checkable. When the content changes, you do not get a fuzzy score, you get a clear "this is not the file that was signed".
You do not need a platform to try this. Hash the output, sign the hash with a key the producer controls, and ship a tiny manifest with it.
import hashlib, json, time
from nacl.signing import SigningKey # pip install pynacl
key = SigningKey.generate() # keep this secret, publish key.verify_key
content = open("report.txt", "rb").read()
manifest = {
"sha256": hashlib.sha256(content).hexdigest(),
"producer": "summarizer-v3",
"model": "model-name-and-version",
"created": int(time.time()),
}
payload = json.dumps(manifest, sort_keys=True).encode()
signature = key.sign(payload).signature.hex()
Verification is the reverse: recompute the hash of the file you received, rebuild the manifest bytes, and check the signature with the public key. If one character changed, the hash will not match.
Provenance does not tell you whether content is true, only who vouches for it and whether it changed. Unsigned content stays unknown, so adoption matters. And key management is the real work: a leaked signing key lets anyone sign anything, so rotate keys and keep them out of app code.
Standards like C2PA already define richer manifests for media. But the core habit is simple enough to start today: sign at the source, verify at the edge, and treat "unsigned" as a fact rather than an accusation.
Would you rather verify where content came from, or keep trying to detect what made it? I am curious which one people think scales.
I wrote a longer, free paper on verifiable claims for public and AI systems, if you want the deeper version: Proof, not promises.