{"slug": "living-models-pairs-gemma-4-with-botanic-1-to-help-decode-plant-dna", "title": "Living Models pairs Gemma 4 with BOTANIC-1 to help decode plant DNA", "summary": "Paris-based AI biology lab Living Models paired Google's Gemma 4 E4B with BOTANIC-1, a plant genomic foundation model pre-trained across 320 plant species, to compress crop genetics research that traditionally takes years of field trials into hours of computation. Gemma 4 E4B runs locally via Ollama on a single NVIDIA L4 GPU as an agentic layer that translates natural-language biological queries into structured queries against BOTANIC-1's inference endpoints, and Living Models found larger models added operational overhead without improving accuracy. The modular pipeline is aimed at helping agricultural scientists narrow tens of thousands of candidate mutations from Bulk Segregant Analysis down to high-confidence causal variants.", "body_md": "# Living Models pairs Gemma 4 with BOTANIC-1 to help decode plant DNA\n\nBy combining Gemma 4 with BOTANIC-1, a foundation model trained specifically on plant genomes, researchers could compress years of traditional crop genetics into hours of computation.\n\nModern crop breeding hinges on pinpointing causation: discovering why one row of melon vines collapses during drought while neighboring vines thrive, or how a single genetic switch alters floral architecture to multiply commercial yield. Tracking that flip across tens of thousands of candidate mutations has historically demanded years of field trials, seasonal plant crosses, and labor-intensive greenhouse screening.\n\nTo speed up this process, Paris-based AI biology lab [Living Models](https://www.livingmodels.ai/) tested a modular architecture pairing Gemma 4 with [BOTANIC-1](https://huggingface.co/collections/living-models/botanic1), a dedicated plant genomic world model. This pipeline demonstrates how modular architecture can untangle decades-old genetic puzzles in an afternoon and help agricultural scientists focus directly on high-confidence interventions.\n\n## Teaching Gemma to read DNA, without teaching it DNA\n\nTo isolate high-value traits, plant geneticists traditionally rely on [Bulk Segregant Analysis](https://en.wikipedia.org/wiki/Bulked_segregant_analysis)—crossing opposing parent lines, cultivating the offspring across a growing season, sorting plants into trait-bearing pools, and sequencing their genomes. While this process narrows the trait to a broad chromosomal neighborhood, it leaves scientists with an unannotated spreadsheet containing tens of thousands of candidate mutations, including letter swaps (Single Nucleotide Polymorphisms, or SNPs), insertions, and deletions.\n\nFaced with this bottleneck, a natural question arises: why not simply train a generalist model like Gemma directly on raw DNA?\n\nDeveloping reliable genomic capabilities within a generalist LLM introduces trade-offs: text-oriented tokenization may be inefficient for nucleotide sequences, supporting long inputs does not inherently capture long-range regulatory interactions, and genomic fine-tuning risks eroding the model’s core reasoning, coding, and tool-use faculties. Conversely, dedicated genomic language models (GLMs) like BOTANIC-1 excel at sequence representations but are difficult for non-technical researchers to deploy, query, and integrate into daily lab workflows.\n\nLiving Models resolves this by pairing the two: Gemma 4 serves as a conversational and agentic abstraction layer, translating natural-language biological inquiries into structured queries against BOTANIC-1’s inference endpoints.\n\n[Gemma 4 E4B](https://huggingface.co/google/gemma-4-E4B) acts as the autonomous bioinformatic orchestrator—writing terminal pipelines and executing tool calls—while BOTANIC-1, a foundation model pre-trained across 320 plant species, evaluates deep evolutionary constraints. Benchmarking showed larger models added operational overhead without improving accuracy, so Living Models runs E4B locally via Ollama on a single NVIDIA L4 GPU. This keeps proprietary genomes sovereign while the minimal footprint slashes operating costs, making domain-specialist AI accessible on standard commercial hardware.\n\n## Beyond correlation: The search for a causal signal\n\nDuring plant reproduction, DNA is inherited in large, contiguous blocks. Harmless mutations hitchhike alongside the causal driver like passengers sharing the same train car. Because every variant in that block co-segregates identically across the sampled plants, statistical sorting software cannot distinguish which letter causes the trait and which are simply along for the ride.\n\nBreaking these genetic ties can be time-consuming. Biologists might cultivate thousands of additional plants to capture rare recombination events, test other plant lines, or spend a year validating individual CRISPR constructs one by one.\n\nTo evaluate where AI creates an actual breakthrough in the lab, the Living Models team tested whether Gemma 4’s agentic reasoning could untangle a trait on its own using:\n\n- **Classical Tools (Setup A):** Deploying Gemma as an autonomous bioinformatician using its agentic reasoning to write pipelines, filter coverage, and execute standard command-line software (bcftools, Ensembl VEP).\n- **BOTANIC-1 (Setup B):** Equipping Gemma with BOTANIC-1, a domain-specific foundation model trained to read the evolutionary grammar of plant genomes.\n\nThe team turned to a landmark discovery led by collaborator Dr. Adnane Boualem, Research Director at INRAE, where a single nucleotide letter swap inside a melon’s *CmEIN3* gene shifts flower sex from female to hermaphroditic—a decisive driver of pollination efficiency and fruit yield. Standard mapping narrowed the field from 47,492 variations genome-wide down to 3,061 regional candidates on Chromosome 2. Yet even on Chromosome 2, classical genetics hits an impassable physical wall: linkage disequilibrium.\n\n### Experiment A: Classical bioinformatics utilities hit a ceiling\n\nEquipped with standard command-line tools, Gemma 4 operated with the procedural discipline of an autonomous computational biologist: formulating an execution plan, applying read-depth filters, and annotating mutations.\n\nThe model exhibited agentic resilience: when syntax errors occurred (such as attempting to apply human-genome flags to plant annotation files) Gemma parsed the terminal outputs and recovered in 9 out of 16 instances. Guided by filtration thresholds, in its best runs, Gemma narrowed the pool to two identical-scoring coding mutations. One was the validated *CmEIN3* driver; the other was an unrelated gene.\n\nTo prove this deadlock was an inherent limitation of classical tools rather than the AI, the team executed the same data through deterministic scripts without the model:\n\nWhile Gemma 4 successfully replicated the ceiling of human expert bioinformatics, that ceiling still leaves researchers at a standstill: because classical tools measure inheritance correlation rather than biological consequence, they cannot identify which mutation actually alters plant function—leaving breeders to spend months of lab validation on a coin toss.\n\n### Experiment B: BOTANIC-1 breaks the tie\n\nIn the second scenario, Gemma 4 was given a single specialized tool: the local BOTANIC-1 scoring function. Instead of relying on population frequency deltas, BOTANIC-1 evaluates variants across up to approximately 128 kb of sequence context using the Log-Likelihood Ratio (LLR) between the probabilities the model gives to the variant (alt) and to the original nucleotide in the reference genome (ref), given the surrounding genomic context (ctx) ([Benegas, G. et al., PNAS, 2023](https://www.pnas.org/doi/10.1073/pnas.2311219120)):\n\nThis metric functions as an evolutionary constraint score; when a nucleotide position has remained conserved across millions of years of plant evolution, introducing an alternative base yields an intensely negative score indicating severe biological disruption.\n\nGemma 4 systematically parsed the Chromosome 2 candidates, filtered out 567 filtered out 567 insertions, deletions and rearrangements, and routed the remaining 2,494 single nucleotide mutations to the local BOTANIC-1 scoring function. The outcome broke the deadlock: the validated *CmEIN3* mutation ranked #1 out of 2,494 candidates without ties.\n\nBy substituting population correlation with deep evolutionary constraint, BOTANIC-1 supplies the orthogonal biophysical signal needed to break through linkage disequilibrium. Equipping Gemma 4 with BOTANIC-1 raised Recall@1 from 0.00 (unguided) and 0.15 (expert‑guided) to 0.90 with no fabricated variants in 20 runs (against 3 and 4 of 20 with classical tools).\n\nThis capacity to shatter genetic ties extends far beyond a single melon chromosome. As detailed in Living Models’ [research](https://www.biorxiv.org/content/10.64898/2026.09.04.749355v1?utm_source=gemini), the team evaluated BOTANIC-1 against classical bioinformatic tools and leading biological models—including PlantCAD2 and Evo2—across more than 500 published, experimentally validated causal plant mutations. Across this diverse retrospective benchmark, BOTANIC-1 placed the true causal mutation in the top 0.1% of candidates in 15.9% of cases, compared to just 4.6% for the best non-GLM baseline, and ranked the causal driver within the top 1.0% in 48.8% of cases (versus 33.9% for classical pipelines).\n\nBy grounding its analysis in deep evolutionary conservation, BOTANIC-1 can equip reasoning models with an empirical understanding of plant genomes that transcends the limits of raw correlation.\n\n## A blueprint for scientific AI\n\nThis modular approach developed by Living Models establishes a repeatable blueprint for scientific AI. Rather than expecting a single model to master every layer of science, modern discovery benefits from a clean separation of concerns: generalist reasoning models like Gemma 4 excel at cognitive orchestration while specialized foundation models like BOTANIC-1 encode the physical and evolutionary constraints of nature. With lightweight models, agricultural labs and developers can execute these agentic pipelines directly on-premises, so intellectual property remains entirely sovereign and secure.\n\nGenetics narrowed our search to a genomic region, but dozens of variants co-segregated identically with the trait. Separating them traditionally means growing and screening thousands of plants over months or years to capture rare recombination events. The real opportunity of AI is dramatically reducing that search space—identifying which constructs to validate first.\n\nAs accelerating climate shifts disrupt traditional growing seasons and challenge global food security, crop breeding can no longer afford to move at the pace of multi-year trial and error. By uniting open reasoning models with specialized genomic intelligence, researchers gain the speed, precision, and privacy needed to engineer resilient crops on a timeline the planet demands.", "url": "https://wpnews.pro/news/living-models-pairs-gemma-4-with-botanic-1-to-help-decode-plant-dna", "canonical_source": "https://deepmind.google/models/gemma/gemmaverse/living-models/", "published_at": "2026-10-09 14:04:09+00:00", "updated_at": "2026-10-09 14:24:56.735318+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-agents", "ai-tools"], "entities": ["Living Models", "Gemma 4", "BOTANIC-1", "Google", "Ollama", "NVIDIA L4 GPU", "Hugging Face"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/living-models-pairs-gemma-4-with-botanic-1-to-help-decode-plant-dna", "markdown": "https://wpnews.pro/news/living-models-pairs-gemma-4-with-botanic-1-to-help-decode-plant-dna.md", "text": "https://wpnews.pro/news/living-models-pairs-gemma-4-with-botanic-1-to-help-decode-plant-dna.txt", "jsonld": "https://wpnews.pro/news/living-models-pairs-gemma-4-with-botanic-1-to-help-decode-plant-dna.jsonld"}}