On AI, human biology, and what it will actually take to discover transformative new medicines
America | Tech | Opinion | Culture | Charts
Today, we’re excited to share a guest post from Daphne Koller, Founder and CEO of insitro. For more from her, subscribe to insitro’s Substack here.
The tech world has latched onto an intoxicating promise: build a superintelligence, and it will cure cancer, and every other disease as well. The logic is seductive. The human body is a system we can already read from and write to, so like other knowledge problems, a powerful enough AI should be able to solve disease.
I fully believe that AI will eventually transform human health. It is why I’ve spent close to 30 years working at the intersection of AI and biology, and the last decade in drug discovery. The need is staggering: by most counts only a quarter of diseases — and by some estimates a few percent — have an approved therapy, and most of those merely slow a disease rather than stop it. For the majority of human illness, medicine still has little to offer.
But the magic-wand promise rests on an assumption that turns out to be false: that we already understand human biology well enough for a clever enough reasoner to find the cures hidden in what we know. We don’t. Hundreds of years into modern medicine, our understanding of most human disease, and much of healthy physiology, is best captured by the parable of the blind men and the elephant; in this case, a really huge elephant. AI is undoubtedly extraordinary, but aimed at a biology we have only begun to measure and barely understand, it will mostly help us generate failures faster.
This essay is about what that actually takes. The path to real cures starts at the root of the problem: measuring biology in the right way and using AI to derive novel insights from those measurements. Below, I describe some of the most prevalent AI Magic Wand narratives, explain the key pitfalls, and offer my perspective on what we actually need to deliver on this important goal.
Three Problems, One Bottleneck
To understand where AI fits, it helps to decompose drug discovery into its three essential stages:
Disease-to-mechanism: Identifying a biological mechanism — a pathway, a target, a molecular interaction — where therapeutic intervention will alter the course of disease in humans.Mechanism-to-drug: Creating a molecular intervention in the right therapeutic modality — a small molecule, antibody, siRNA, gene therapy — that achieves the desired mechanistic effect with acceptable safety and pharmacological properties.Drug-to-patient: Designing a clinical development program that identifies the right patients and assesses the molecule’s effects — beneficial as well as adverse.
The vast majority of AI work in drug discovery has focused on stage 2. This is understandable: the origin of the AI Magic Wand exuberance is the incredible achievement of AlphaFold, a field-defining tour de force. From this starting point, we have seen an explosion of AI tools capable of designing novel proteins, small molecules, RNA therapies, and even gene therapies. Given a biological mechanism we want to hit, it seems that AI can now design a molecule to hit it faster and better than ever before.
AI will certainly generate new and better molecules at an unprecedented rate, but will it generate drugs that unlock diseases for which there is currently no meaningful treatment? There are “undruggable targets” — high-confidence mechanisms that historically we have been unable to hit. The success against KRAS, the quintessential undruggable target, shows that this journey is possible. Notably, this success emerged from decades of structural biology and medicinal chemistry, not AI; as of now, I don’t know of a single example of an AI-derived insight that has led to “drugging the undruggable.” More broadly, the biggest step functions in our ability to drug the undruggable have historically come not from better molecular design tools, but from expanding our repertoire of therapeutic modalities: first biologics, then siRNA and antisense oligonucleotides, then gene editing. Each new modality opened a class of targets that was simply inaccessible before.
But an even more critical point: validated yet undruggable targets are a tiny handful in the landscape of unmet need. For the vast majority of diseases without effective treatments, we simply have no idea what the right mechanism is. More than 90% of drugs that enter clinical trials fail — a dismal statistic that has barely improved in several decades. In the large majority of cases, the molecule was engineered just fine. The mechanism it targeted was wrong. We are doing a pretty good job at manufacturing keys, but they are generally for the wrong locks. Even if AI lets us make better keys at an accelerating pace, that won’t improve our ability to identify the right locks. **The real bottleneck in making a novel medicine is disease understanding: identifying a biological mechanism whose modification actually changes the course of disease in patients. **That, far more than molecular design, is where drug discovery succeeds or fails.
This mechanistic understanding is a rare commodity. And because no one likes to fail in the clinic, we are seeing industry trends that are truly destructive. There are currently 38 targets that have over 50 programs against each of them — slightly better keys for those few locks where we have strong conviction. How many variants of GLP-1 do we really need? Even worse than this misallocation of capital is the disservice to patients: the number of novel targets the industry advances each year fell from ~100 in 2015 to about 30 in 2024. That collapse is the far bigger cost: the inability to help the hundreds of millions of people for whom medicine currently offers nothing.
The Data Chasm of Human Biology
The disease-understanding goal requires that we bridge the chasm between high-level clinical manifestations of disease in a patient and the granular cellular mechanisms in which a drug intervenes. This has given rise to a second manifestation of the AI Magic Wand. Large language models — with their super-human reasoning capabilities — will connect the dots across the vast published literature of human biology, reasoning their way to new mechanistic hypotheses.
This optimism makes a very strong assumption: that the scientific community has collected — or will soon collect — enough data about human biology to contain the answer, and that we just need better reasoning to extract it. The challenge is that human biology is incredibly complex, spanning multiple interconnected biological layers — DNA, protein, cells, multi-cellular environments, entire organisms. Individual components respond dynamically to even subtle changes in related components or in the environment. Moreover, biology wasn’t engineered; it is the result of billions of years of messy, stochastic evolution, which produced staggering variation — countless genes, cell types, states, and contexts, each behaving in its own way. There is too much of it, too idiosyncratic, to reason about in the abstract. You have to measure it.
How many experiments would we need? The table below is an informed attempt at a back-of-the-envelope analysis.
Even if we consider only cell biology — the layer we need to interrogate biological mechanism — the space is vast. It becomes exponentially more vast when we consider that a drug is an intervention, so we need to map not only biology as it is, but also how it would respond to a perturbation. The largest cell atlases assembled to date, now spanning hundreds of millions of cells, remain orders of magnitude too small to cover this space. Until recently they also held almost no causal, perturbational data, the measurements most critical for understanding what an intervention would do in a living system. That has begun to change: several organizations have launched the monumental effort of building a “Virtual Cell,” pairing large-scale perturbation data with AI to reduce the data collection burden. But even the largest of these efforts samples only a vanishing fraction of the possible perturbations, and does so almost entirely in a narrow range of cell lines.
Even more importantly, the virtual cell efforts — useful as they might eventually turn out to be — do not address the other half of the equation: relating biological mechanisms to human clinical outcomes. Most human disease is a systems-level dysfunction, involving a complex, temporal interaction of multiple biologies spanning diverse cell types. Understanding these processes requires systems-level measurements that are far less scalable, often requiring living organisms.
Which brings up the greatest data challenge. While some processes are conserved across all forms of life, others are far more specific. The folding of a single protein is a self-contained process, highly conserved — closer to physics than to biology; this allows protein folding models to be trained on sequences collected across thousands of species. Metabolism involves at least a dozen distinct cell types and might be conserved across mammals. Brain function and dysfunction involves dozens of distinct cellular identities; and these processes are exquisitely specialized to humans: rodents do not get Alzheimer’s disease; non-human primates do not recapitulate ALS. The diseases where we have made the least progress tend to be precisely those that are most human-specific, and therefore those for which the data is most expensive to collect, least available, and most fraught with ethical constraints.
The Right Experiments, Not All Experiments
But we don’t need a universal causal model in order to make medicines. We can instead build a targeted causal model that makes the space navigable toward a desired outcome — uncovering the biological mechanisms underlying human diseases. Our models will not cover all biologies or all diseases, but an informed selection process will still allow us to make considerable headway.
This insight motivates the thesis behind a third AI Magic Wand: swarms of AI agents operating in a closed loop with laboratory automation, relentlessly chipping away at scientific problems. They formulate hypotheses, direct robots, analyze results, and iterate. Early successes here have been beguiling — automated systems have already excelled at tasks like optimizing cell-free protein synthesis or designing antibodies against known targets.
But agentic iterated optimization relies on a fundamental attribute: agents thrive when there is a fast, accurate, and cheap scorecard to evaluate progress. If you give a sufficiently smart model an instant feedback loop, it will grind against that benchmark until it wins. This is why coding assistants and molecular design tools advanced so rapidly — the feedback is cheap, accurate, and fast. A compiler immediately verifies whether code will run. The iterative loop with a developer provides rapid feedback on intent. The closed loop is tight, cheap, and objective.
Drug development is the exact opposite. The ultimate scorecard — whether a drug actually provides therapeutic benefit to a patient — cannot be captured well by computational models or high-throughput assays. The only true ground truth is a human clinical trial. This feedback loop currently takes years, costs millions, and is strictly bound by human ethics and living biology. It is the ultimate slow feedback loop, and no amount of compute or process optimization can change this.
The agentic lab successes listed above are solving problems that look like coding: highly quantitative objectives within a constrained search space. These problems are valuable, but largely beside the point when it comes to predicting whether a drug will actually work in a patient population. You cannot solve the human translation problem by accelerating our ability to optimize the wrong objective function. Doing so will simply scale up our process for building more keys to the wrong locks.
In order to leverage the power of agentic AI and lab automation towards the goal of improving scientific discovery, we must first have an objective function that is a proxy to human clinical benefit, and yet allows for AI insights and rapid experimentation.
The Final Mile: Patient Impact
Some have argued that the most important AI unlock in drug discovery is in the third stage — drug-to-patient — taking a drug candidate through preclinical testing and clinical trials. This is the fourth AI Magic Wand: reduce the time and cost of this very expensive phase, and drug discovery becomes faster and cheaper. Sadly, if you accelerate a pipeline full of drugs aimed at the wrong mechanisms, all you get is faster failures.
At the same time, there are real opportunities for AI in clinical development: predictive toxicology models and AI-drafted regulatory filings can shorten IND-enabling work, and patient identification from electronic health records, smarter site selection, and automated data management can trim operational overhead. These compression levers offer meaningful benefits, and we should absolutely pursue those. But these gains sit almost entirely within the first two slices of work. The remaining 50% — in-life biological observation — is gated by how fast disease unfolds in a living human. Even a comprehensive AI-driven improvement across everything it can touch leaves the majority of development time and cost structurally unchanged.
There is, however, one pathway through which AI could accelerate even this irreducible clock — and it runs directly through biological mechanism. An AI-enabled, deep mechanistic understanding of a disease enables the identification of novel clinical readouts that serve three distinct purposes: selecting the patients most likely to respond, confirming that the drug is hitting its intended target, and detecting early and reliable signals that it is actually modifying disease biology. Together, these allow trials to enroll the right patients, read out faster, and catch failures earlier — changes that transcend clinical trial operations, transforming the trial design itself. This capability is inseparable from solving the disease-understanding problem; they are one and the same. Better trials, in the end, are downstream of better biology.
Fulfilling AI’s Promise
The AI Magic Wands described above have real value. Generating molecules quickly can be a significant accelerant, as is reasoning across the vast scientific literature and increasing the efficiency of the scientific process. But these tools are just that: point solutions that address important problems without altering the fundamentals of our industry.
To fulfill the promise of AI for the millions of patients lacking any meaningful treatment, we must direct our efforts toward the problem that really matters: the identification of biological mechanisms with disease-transforming clinical benefit. This is arguably the hardest problem in drug discovery, because the only conclusive test of whether we have correctly identified a novel biological mechanism is a human clinical trial. There are multiple other paths in this space with shorter timelines and clearer near-term proof points. Those paths are shorter because the problems are more tractable: the feedback loops are faster and the benchmarks are cleaner. But a shorter path to a smaller destination is still a smaller destination — process improvements for problems we already know how to solve.
For the hundreds of millions of patients for whom no meaningful medicines exist, the difference between a wrong mechanism and the right one is the difference between another devastating clinical failure and a life-altering breakthrough. That is what we need in order to truly deliver on AI’s promise for human health. This harder but aspirational path is the one we have elected to take at insitro, and we’ll have more to say about our approach soon.
This newsletter is provided for informational purposes only, and should not be relied upon as legal, business, investment, or tax advice. Furthermore, this content is not investment advice, nor is it intended for use by any investors or prospective investors in any a16z funds. This newsletter may link to other websites or contain other information obtained from third-party sources - a16z has not independently verified nor makes any representations about the current or enduring accuracy of such information. If this content includes third-party advertisements, a16z has not reviewed such advertisements and does not endorse any advertising content or related companies contained therein. Any investments or portfolio companies mentioned, referred to, or described are not representative of all investments in vehicles managed by a16z; visit https://a16z.com/investment-list/ for a full list of investments. Other important information can be found at a16z.com/disclosures. You’re receiving this newsletter since you opted in earlier; if you would like to opt out of future newsletters you may unsubscribe immediately.