Duke’s Raygun AI Shrinks Proteins to Unlock Gene Therapy Duke University researchers have developed Raygun, an AI tool that shrinks proteins while preserving function, potentially easing gene therapy delivery. The system generated 70,000 fluorescent protein variants, with the smallest functional ones at 199 and 206 amino acids, shorter than 96% of known fluorescent proteins. Published in Nature on July 29, Raygun could reduce costs for therapies like Hemgenix ($3.5 million per dose) and Zolgensma ($2.1 million). July 31, 2026 , Inside AI — A new artificial intelligence system can shrink proteins to a fraction of their natural size while preserving their function, a capability that could crack open bottlenecks in gene therapy, drug delivery, and basic research. Scientists at Duke University have developed Raygun , an AI tool that compresses, enlarges, or rewrites protein sequences using a combination of protein language models and a novel fixed-resolution encoding scheme. The work was published in Nature on July 29 . The team demonstrated Raygun’s power by generating 70,000 variants of two fluorescent proteins. After computational filtering, eight were synthesized in human cells, and six glowed. The smallest working variants were just 199 and 206 amino acids long, shorter than 96% of all known fluorescent proteins in standard databases. Gene therapy faces a persistent size limit. Therapeutic genes must be packed into hollowed viral shells, but many disease-correcting genes are too large to fit. Shrinking the protein cargo could enable delivery with cheaper, better-characterized vectors. Hemgenix for hemophilia costs $3.5 million per dose. Zolgensma for spinal muscular atrophy costs $2.1 million . In India , families often resort to crowdfunding for a single treatment. “Design tools that shrink a therapeutic protein without breaking it widen what can be delivered by cheaper, better-characterised vehicles,” said Rohit Singh , assistant professor of biostatistics and bioinformatics at Duke and co-author of the study. Raygun’s approach rests on two innovations. First, it summarizes every protein at a fixed resolution regardless of length. Kapil Devkota , the study’s lead computational author, divided each sequence into the same number of blocks, averaging within each block to create a uniform-length representation. This lets the model compare and convert between proteins of different sizes. Second, Raygun treats a protein not as a single sequence but as a probability distribution over plausible variants. It can sample a new sequence at any target length in roughly 0.3 seconds on a single graphics processor, about 100 times faster than diffusion-based methods. The system learns from evolution. Protein language models, trained on hundreds of millions of sequences, capture which parts of a protein are conserved across species and which tolerate change. Raygun then re-expresses the protein’s core function in fewer amino acids, guided by two user controls: how far to stray from the original and how long the result should be. The fluorescent protein results hint at deeper editing capabilities. Some working variants carried more than 40 coordinated insertions, deletions, and substitutions. Uncoordinated edits, by contrast, destroy fluorescence after just 4 to 5 changes. But the engineered proteins glow dimly. Reaching the brightness of laboratory standards will require further rounds of directed evolution. A more fundamental limitation stems from Raygun’s reliance on natural sequences: it preserves what evolution selected for, not what human engineers later added. “Which properties of an engineered protein survive an AI edit, and which do not, is an open question that we and other researchers are now actively tackling,” Singh said. The work also raises dual-use concerns. The same model that designs a novel protein could assess whether an unknown design resembles something dangerous. The team are signatories to the Responsible AI for Biodesign principles. Raygun’s speed and controllability could accelerate protein engineering beyond gene therapy, from industrial enzymes to biosensors. The underlying method of fixed-resolution encoding may also influence how other biological sequences are modeled. The study appears alongside broader efforts to apply language models to protein design, including work from Meta ’s ESM team and DeepMind ’s AlphaFold successors. Raygun’s focus on length changes addresses a gap that structure-prediction tools alone cannot fill. The code and model weights have been released for academic use.