OpenADMET: Response to Pat Walters Peter Kenny responded to Pat Walters' critique of the OpenADMET initiative, defending the role of machine learning in modern drug discovery. Kenny argued that Walters mischaracterized the utility of ML models, noting that they actively drive most discovery programs, while Walters maintained that clinical decisions rely on measured assays and that universal ML models for ADME properties are unlikely. Pat Walters responded https://patwalters.github.io/Response-to-Peter-Kenny/ to my post https://fbdd-lit.blogspot.com/2026/07/the-openadmet-initiative.html on the OpenADMET https://openadmet.org/ initiative and I am responding to his post. As is normal for posts here my responses are italicised in red and enclosed in square brackets. Many thanks, Pat for taking the time to respond to my comments. A recent post on the Molecular Design blog by Peter Kenny critiques the OpenADMET initiative. While constructive debate is essential for open science, the post contains several fundamental misconceptions about the roles of machine learning, structural biology, and ADMET optimization in modern drug discovery that need clarification. Peter, While I appreciate your advocacy for Open Science and your pragmatic view on the complexity of human biology, your critique of the OpenADMET initiative relies on several fundamental mischaracterizations of how modern computational chemistry operates and what the initiative actually aims to achieve. Here are the key areas where your arguments fall short: 1. Dismissing ML Model Utility Ignores Current Industry Reality You argue that project teams historically delivered clinical candidates without predictive models, and suggest that we need “new assays that are more predictive… as opposed to new ML models.” This severely underestimates how drug discovery is actually conducted today. This is a misrepresentation of what I argued although I concede that the points could have been more clearly articulated . I'm certainly not dismissing ML modelling in drug discovery although I remain sceptical that it will prove to be the panacea that some seem to think that it will be. First, I actually stated that “drug discovery project teams have delivered and continue to deliver clinical development candidates without ever having sufficient data for building ML models that can accurately predict all the quantities of interest to the project teams”. What I was getting is that some decisions that shape the future of drug discovery projects get made before there is enough project-specific data to enable ML models to be built. This will be less of an issue if models for predicting affinity such as Boltz-2 do indeed turn out to be universal. However, I consider it less likely that universal models will be found for prediction of ADME-related properties such as aqueous solubility, permeability and turnover by CYPs. Second, I stated “All that said, decisions to take compounds into clinical development are based on measurements made in a range of assays and failure in clinical development reflects an inability of these assays to predict clinical outcomes. To more effectively address attrition we actually need new assays that are more predictive of outcomes in clinical development as opposed to new ML models that are more predictive of quantities that will need to be measured anyway.” This does not constitute dismissal of ML model utility and I'm merely making the point that decisions as to whether to take a compound into clinical development are based on assay measurement as opposed to predicted assay results. I certainly believe that better predictions for ADME and off-target bioactivity will lead to faster discovery of clinical development candidates. However, clinical outcomes are uncertain even when you've got a full set of measured data bear in mind that PK/PD modelling to set doses for Phase 2 trials uses measured human PK data from Phase 1 . ML models actively drive modern discovery: The vast majority of drug discovery programs across pharma, biotech, and academia are now directly driven by ML models. While I would not dispute that ML models are widely used in drug discovery, your claim that ML models are actually driving the majority of drug discovery programs is in need of evidence. Far from being academic novelties, predictive models are integrated into daily multi-parameter optimization MPO cycles to guide design, prioritize synthesis, and shorten cycle times before a molecule ever touches a wet-lab assay. Models complement, rather than replace, assays: Models do not eliminate the need for physical assays; they ensure that expensive synthesis and wet-lab capacity are spent on the most promising molecules. Relying purely on physical assays for every ideated variant is wildly inefficient. Agreed and I never suggested otherwise. I was simply making the point that physical assays are required to generate the data for selection of candidates for clinical development. The throughput mismatch requires models: While ADMET modeling typically focuses on local lead-optimization spaces rather than vast, multi-billion-compound virtual libraries, the number of ideated compounds evaluated in any series is still orders of magnitude larger than the number that can realistically be synthesized and assayed. Predictive models are essential to filter that ideation space down to the highest-probability candidates. Agreed and I never suggested otherwise. I was simply making the point that physical assays are required to generate the data for selection of candidates for clinical development. Expanding and improving the drug discovery toolbox: Assuming that status quo methods are sufficient ignores the stark reality that drug discovery is more expensive and slower than ever Eroom’s law https://en.wikipedia.org/wiki/Eroom%27s law and faces intense economic and productivity scrutiny across the entire biopharma sector. Building predictive models is about continuously expanding and refining the drug discovery toolbox, equipping teams with better decision-making capabilities to accelerate pipelines and lower skyrocketing costs. Agreed and I never suggested otherwise. I was simply making the point that physical assays are required to generate the data for selection of candidates for clinical development. 2. Semantic Gatekeeping vs. Structural Reality A significant portion of your critique focuses on terminology “ADMET projects a lack of familiarity”, “Avoid-ome”, “rational drug design” . This semantic distraction misses the underlying biophysical reality: The underlying biophysical reality is that binding to a CYP and being turned over by a CYP are very different phenomena. Unified biophysical mechanism: ADME clearance drivers e.g., CYP metabolism, P-glycoprotein efflux and off-target toxicities e.g., hERG inhibition, anti-target binding are fundamentally unwanted protein-ligand interactions . While I generally agree that active efflux is fundamentally unwanted I would argue that CYP metabolism is actually essential although it certainly needs to be carefully controlled. For example, it might be desirable for the drug to be turned over by two different CYPs to reduce the potential for drug-drug interactions. My view is that turnover of of compounds by CYPs shouldn't be described as "protein-ligand interactions" even though turnover starts with a non-covalently bound complex. It is a functional catch-all, not an absolute mandate: Nobody is suggesting that “Avoid-ome” targets must be completely avoided in a literal sense. In your Nature Communications article you are stating that the Avoid-ome proteins are anti-targets Fig. 1: The set of protein anti-targets that comprise the Avoid-ome and many drug discovery scientists would take the 'anti-target' label to mean a protein that must be avoided. You're also extending the usual definition of an anti-target from a protein that is associated with toxicity risk when engaged in vivo to also cover metabolic enzymes and transporters you're allowed to do this but you do need to let your readers know . Drugs must be metabolized, and binding to abundant proteins like Human Serum Albumin HSA is a fundamental aspect of pharmacokinetics. The “Avoid-ome” is simply a convenient catch-all term for the network of proteins governing clearance, distribution, and toxicity. Framing these interactions within a unified, structure-based computational framework makes complete sense. I certainly agree that it would be beneficial to bring diverse phenomena such as reversible binding and turnover by catalytic enzymes into a unified, structure-based computational framework. However, there is rather more to achieving this than simply declaring that proteins such as hERG, HSA and CYPs are all Avoid-ome anti-targets. Focusing on gatekeeping over substance: Claiming that “using the term ADMET projects a lack of familiarity with the practical realities of drug design” is an ad hominem distraction. ADMET is a well-understood term widely used and understood by experienced practitioners. Quibbling over whether an acronym is traditionally split or grouped does nothing to change the underlying science: are off-target binding, clearance, and toxicity critical to model and optimize, or not? I used "projects" so as to avoid ad hominem issues and the context in which the term ADMET is used is important. There are a number of points that I think your Nature Communications article needed to make and I must stress that my criticism is that you and your co-authors have have failed to acknowledge issues and not that you don't understand the issues . First, you needed to demonstrate awareness of the distinction between pharmacodynamics and pharmacokinetics. Second, you needed to acknowledge that different phenomena such as reversible binding and turnover generally have to be modelled differently. Third, you needed to state that you were using the normal definition for anti-targets. 3. SBDD Applies to Off-Targets Just as It Does to Targets of Interest Your critique downplays the utility of structural data for ADMET, but structural information is absolutely key to understanding and modulating the interactions between drugs and off-target proteins. It's not accurate to state that I downplayed the utility of structural data in my post and I actually stated: "I see a degree of overlap between the OpenADMET and OpenBind initiatives in that safety assessment will often require prediction of binding affinity of anti-targets for compounds being considered for synthesis. Indeed, there is no reason that structures for complexes of anti-targets with ligands should be excluded from data sets when the objective is to build universal ML models for prediction of binding affinity." In my post on OpenBind I stated: "Given the focus on enabling affinity prediction, there is no reason for excluding anti-targets or non-human proteins." That said, the usual basis for SBDD is reversible binding and you need to escape from the constraints of the reversible binding mindset if you want to present credible solutions to ADME issues such as turnover by metabolic enzymes. - Extending Structure-based drug design SBDD to anti-targets: SBDD principles are not limited to primary therapeutic targets. Having high-resolution structural information for off-targets, transporters, and metabolic enzymes allows medicinal and computational chemists to apply the exact same rational SBDD strategies such as optimizing steric clashes or altering hydrogen-bonding patterns to dial out unwanted off-target binding or tune metabolic stability. I stated in my post that "with respect to to 'atomistic detail' it's important to bear in mind that structures for transition states relevant when the quantity of interest is rate of turnover by metabolic enzymes cannot actually be observed in experimental protein structural studies". In these scenarios invoking "ground truth" might come across as arm-waving. Using PXR as an example, a 2021 perspective https://doi.org/10.1021/acs.jmedchem.0c02245 by UCB scientists showed that SBDD approaches added significantly to their understanding of and mitigation of PXR liabilities. Fueled by high-quality open data: The integration of protein structure with machine learning is still a nascent field, but it holds immense promise. Agreed although you do need to concede that docking and scoring have been around for decades and I'll direct you to my post https://fbdd-lit.blogspot.com/2026/06/the-openbind-initiative.html on the OpenBind project and some of the discussion there might be relevant to the use of protein structural information to build ML models for affinity. Providing high-quality, standardized experimental data paired with structural ground truth is precisely what is needed to advance the field and unlock next-generation predictive models. 4. Active Learning Identifies Optimal Experimental Strategies for Model Building You raise concerns about the difficulty of covering chemical space at a fine enough resolution to build generalizable models. My view is that few if any QSAR models were shown to be usefully-predictive outside the chemical space structural series, for example in which they were trained and I would be surprised if OpenADMET can meaningfully sample structural series that have not yet been conceived. I accept that models such as Boltz-2 can should enable prediction of binding affinity for chemotypes that haven't been sampled in training. However, the ability to predict binding affinity is of minimal value in ADME modelling. - Focusing on optimal data generation: OpenADMET does not rely on brute-force, random assay generation. A primary objective of the initiative is specifically to identify optimal strategies for using targeted experiments to build better, more predictive models. - Iterative efficiency: By using active learning, the initiative strategically samples focused chemotypes to maximize model performance with the fewest physical assay points. This directly answers the challenge of sparse data in lead optimization without wasting resources on irrelevant chemical space. 5. Biological Complexity Is Not a Reason to Ignore Predictable Failures Concluding that the primary challenge in drug discovery is “the uncertainty that results from human biology” rather than ADMET optimization presents a false dichotomy and dodges an actionable problem: I asserted that is that "the principal challenge for drug discovery is and has always been the uncertainty that results from the complexity of human biology" in response to your claim: " Understanding and navigating the Avoid-ome is the central universal challenge of modern drug discovery." 30%+ of failures are preventable: Unpredictable efficacy due to complex disease biology is indeed a major hurdle, but roughly 30% of clinical failures still stem from preventable flaws in safety and pharmacokinetic profiles. We cannot solve biological complexity by ignoring predictable off-target interactions. You do need to present evidence in support of this claim and I note that the three references that your article cites in connection with attrition are from 2004, 2015 and 2014 respectively. When you assert that failures are preventable bear in mind that the those who nominate compounds for clinical development will have measurements for all the quantities for which you seek to provide. Better ADMET tools help unpack biological complexity: Understanding ADMET better makes us much faster at reaching clean tool compounds and drug candidates. I fully agree that a better understanding of ADMET will enable clinical development candidates to be identified more quickly. Chemical probes are typically used in cell-based assays for target-validation and in these situations ADME is entirely irrelevant in these situations as are most toxicities you don't need to worry about hERG if the cells that you're expressing doesn't express it . Chemical probes need to be very potent and highly selective even more so than for drugs with respect to proteins that closely related to the target of interest. My understanding is that probes used in vivo tend to be compounds from discovery projects that are no longer under consideration as clinical development candidates I'm not sure how much effort is devoted to design of probes for in vivo studies . Here's some information https://www.chemicalprobes.org/info/what-are-chemical-probes on from the excellent Chemical Probes Portal that I hope you find helpful. These high-quality molecules enable researchers to test biological hypotheses more quickly and reliably in vivo, which is precisely how we address and overcome the challenge of biological complexity, Summary Dismissing predictive ML models overlooks the fact that they are already driving the majority of active drug discovery pipelines globally. OpenADMET’s mission to systematically map off-target structural space with high-quality open data is not an academic exercise: it is the foundational infrastructure modern SBDD and ML-driven drug design require to reduce clinical attrition.