cd /news/artificial-intelligence/new-ai-model-gpn-star-learns-from-ev… · home topics artificial-intelligence article
[ARTICLE · art-124874] src=insideai.news ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

New AI Model GPN-Star Learns from Evolution to Decode Human Genome

UC Berkeley researchers developed GPN-Star, a genomic language model that identifies disease-linked genetic variants using whole-genome alignments instead of raw DNA sequences, slashing training time from months on 2,000 NVIDIA processors to days or hours on a handful. The model, published in Nature, predicts pathogenicity of non-coding variants and flags functional versus non-functional genome elements, with predictions released to help prioritize experiments for human health.

by read3 min views4 publishedSep 9, 2026
New AI Model GPN-Star Learns from Evolution to Decode Human Genome
Image: Insideai (auto-discovered)

September 9, 2026, (Inside AI) — A new genomic language model called GPN-Star identifies disease-linked genetic variants with far less computing power than rivals. Researchers at UC Berkeley trained it on whole-genome alignments instead of raw DNA sequences.

The model predicts which mutations in non-coding DNA affect human health. It also flags functional versus non-functional elements across the genome. The team published genome-wide predictions alongside the study in Nature.

More than two decades after the first full human genome sequence, most of its 3 billion base pairs remain poorly understood. Only 1 to 2% codes for proteins. The rest includes regulatory regions and evolutionary leftovers often called junk DNA.

These non-coding regions likely influence cancer, heart disease, autism, and other inherited conditions. But scientists cannot experimentally test every variant. Computational prioritization is essential.

"Our model excels in making predictions about the pathogenicity of genetic variants, and identifying functional versus non-functional elements in the genome," Yun Song, professor of computer science and statistics at Berkeley and investigator at the Innovative Genomics Institute

Song is senior author of the study. He also directs the Berkeley Center for Computational Biology and co-directs the UC Berkeley-UCSF Bakar Computational Biomedicine Initiative.

Whole-genome alignments cut training costs #

Most genomic language models learn from unaligned genomes across many species. The massive Evo 2 model trained on over 100,000 species. It required 2,000 NVIDIA processors and months of compute.

GPN-Star takes a different path. It uses whole-genome alignments that map hundreds of species to a single reference genome. This highlights conserved code and evolutionary changes before training begins.

The approach slashes training time to days or even hours on a handful of processors. It also reduces confusion from junk DNA that dominates most genomes.

"We tried to help the model learn by curating data that's more likely to harbor functional elements," Song

"Our approach is that we should use these biological insights to improve the model, rather than hoping that the model will figure out what's important by itself."

Evolutionary timescales shape prediction strengths #

The team trained GPN-Star on human-anchored alignments at three evolutionary scales: primates, mammals, and vertebrates. They also trained models for mice, fruit flies, chickens, roundworms, and a plant species.

Different timescales optimized different predictions. Longer timescales better predicted rare protein-coding variant impacts. Shorter timescales better predicted complex trait risks like schizophrenia.

"We found that models trained at different evolutionary time scales were actually optimized for interpreting different kinds of genetic variants," Chengzhong Ye, graduate student in statistics at UC Berkeley and study co-first author

Complex traits can involve up to 10,000 mutations, many in non-coding regions. Primate-specific training improved those predictions.

"For complex traits, we were surprised and pleased to see that training a model that's specific to primate genomes, which are more relevant to recent human evolution, really helped us make better predictions," Song

The researchers hope low training costs let other labs adapt the model. They released genome-wide predictions for biologists to prioritize experiments.

"We hope our work will help drive biological discovery," Song

"People have developed really creative tools for assaying the impact of genetic variants, but they cannot experimentally test every single variant in the genome. We believe our predictions will help to prioritize the experiments that could have the greatest impact on human health."

"We're making great progress," Gonzalo Benegas, study co-first author

"But the more people that can work with these models, the better they will get."

Additional co-authors include Carlos Albors, Canal Li, Sebastian Prillo of Berkeley, Peter Fields of Jackson Laboratory, and Brian Clarke of the German Cancer Research Center. The National Institutes of Health and the Chan Zuckerberg Initiative provided funding and GPU resources.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @uc berkeley 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/new-ai-model-gpn-sta…] indexed:0 read:3min 2026-09-09 ·