# Protecting the Authenticity of the Genome Record

> Source: <https://projecttimestamper.org/blog/genomes/>
> Published: 2026-08-24 21:31:34+00:00

We’re thrilled to announce the completion of a new effort to cryptographically timestamp more than 3.6 million genome sequences found in nature! This work preserves the authenticity of the world's public genomic record, including genomes from animals, plants, fungi, bacteria, archaea, viruses, and *Homo sapiens*.

Just as we have timestamped existing great works of humanity, including literature, art, music, film, and science, this project timestamps existing great works of Mother Nature: millions of genomes that together contain the design of our biosphere. There is surely no dataset more important to humanity's understanding of Earth and its living things.

The AI revolution, with its escalating power to generate vast amounts of data, to hack databases, and to design new genes and new organisms, threatens to raise questions that were previously unthinkable. Can large scientific databases be largely fake? Are we looking at real scientific data, or counterfeit? Is an organism found in the wild natural or synthetic? Are a pathogen's genes what we think they are? Is the human genome what we think it is? We felt that timestamping the current dataset of genome sequences was urgent, before the risk of tampering grew.

Our new timestamps cover the [full public collection of assembled genome sequences](https://www.ncbi.nlm.nih.gov/datasets/genome/) from the database maintained by the National Center for Biotechnology Information (NCBI), run by the US National Institutes of Health. These sequences were generated by the work of thousands of researchers in scientific labs around the world over the past couple of decades. First, researchers used laboratory techniques to extract, purify and break up DNA from cells; next, they used machines to run large-scale sequencing of the DNA fragments; finally, they used software to take the resulting sequence fragments and assemble them into genome sequences.

Our script ran for a few weeks to download each individual genome sequence file from NCBI and digest it using a cryptographically secure algorithm (SHA-256). The resulting digests were subsequently collected in hash lists, which were submitted to the OpenTimestamps service for timestamping on the bitcoin blockchain. For our purposes, we don't need to store the sequences themselves; we keep only the hashes and timestamps.

By timestamping these genomes, we have constructed a permanent, verifiable proof that the genome sequences, as currently recorded in the NCBI database, existed as of August 21, 2026. Attackers attempting to modify NCBI’s database to forge new counterfeit genome sequences, or to modify existing genome sequences, will be unable to provide such a proof that the fake sequences existed that early.

We are thereby helping to preserve the authenticity of existing genome sequences in a future when it will be easier and cheaper to generate a simulated genome than to obtain the real sequence from a physical organism. Future historians, biologists, and medical researchers will be able to refer to these datasets with reasonably high confidence that the sequences came from genuine DNA sequencing of real, natural organisms that existed before large-scale database hacking and large-scale bio-hacking were possible.
