The Role of Rigor in Artificial Intelligence A new paper by an AI researcher argues that the common criticism of deep learning as 'alchemy' misses the mark, proposing instead a framework of three forms of rigor—conceptual, epistemic, and operational—to explain AI's progress and uncertainty. The paper highlights that while operational rigor (benchmarks, evaluations, deployment controls) is strong, conceptual and epistemic rigor lag, leaving basic questions about understanding, generalization, and failure unanswered. The author calls for balanced development of all three forms to mature AI as both a science and a technology. “Deep learning is alchemy” may be the most repeated criticism in AI. It also misses the mark. Alchemy failed to deliver results. Deep learning, by contrast, has produced transformative technologies. And fields like medicine are only partially understood without being deemed alchemical. So calling AI “alchemy” captures part of the problem, but not all of it. Modern AI is not simply undisciplined experimentation. It contains significant amounts of rigor. But we still struggle to answer basic questions: • Do models understand? • Why do they generalize? • When will they fail? The deeper issue is that rigor takes different forms—and in AI, those forms are unevenly developed. My new paper distinguishes three: • Conceptual rigor: coherent terminology and paradigms • Epistemic rigor: reliable scientific understanding • Operational rigor: reliable performance and deployment This framework helps explain both the extraordinary progress of modern AI and the uncertainty surrounding it. Conceptual rigor asks whether the field knows what it's talking about. • What exactly is intelligence? • What qualifies as AGI? • What does it mean for a system to be aligned? Consider the debate over whether current models are intelligent. One person points to their breadth of performance. Another points to weak planning. Another emphasizes sample inefficiency. Another asks whether it has a grounded model of the world. They appear to disagree about one property. Often, they are evaluating four. This is why conceptual clarity matters in practice. Questions about intelligence, understanding, AGI, and alignment do not remain confined to philosophy: they shape how things are measured, optimized, and built. Epistemic rigor asks whether empirical success has become scientific understanding. The paper focuses on three criteria: • Can findings be reproduced? • Can behavior be predicted in advance? • Can success and failure be explained? AI experiments are unusually reproducible in principle: code, data, and models can be copied. But conclusions may still depend heavily on random seeds, hyperparameters, implementation choices, benchmark selection, and compute budgets. Reproducing a number is not always the same as reproducing the conclusion drawn from it. Prediction is harder. Scaling laws can forecast some training outcomes. Infinite-width theory can lead to more tractable settings. Classical learning theory explains important pieces. But we still lack broad principles telling us when a model will generalize, fail under distribution shift, or remain robust under adversarial perturbations. Explanation is harder still. Neural networks are mathematically specified, yet their learned features resist human interpretation. A behavior may arise from training data, optimization dynamics, internal representations, or interactions among all of them. The system is transparent in code but opaque in meaning. Operational rigor is where modern AI is strongest: benchmarks, evaluations, monitoring, red-teaming, and deployment controls. The field has become highly effective at improving systems without first obtaining a scientific theory of them. Benchmarks turn capabilities into measurable targets. Post-training shapes behavior. Tools and scaffolding compensate for model weaknesses. Operational rigor can therefore partially substitute for scientific understanding. That imbalance defines the deep-learning era: • Capabilities rise rapidly. • Explanations lag behind. • Benchmarks become optimization targets. • New systems generate new phenomena faster than theory can absorb them. AI is advancing while continually changing the object that science must explain. For AI to mature as both a science and a technology, it will require all three forms of rigor: • Clearer concepts to define our goals. • Stronger science to predict and explain system behavior. • Better engineering to make systems genuinely reliable. The future of AI depends not simply on demanding “more rigor,” but on identifying which kind is missing—and understanding how the imbalance shapes what we can build, know, and control. - Paper link: arxiv.org/abs/2607.03634 https://arxiv.org/abs/2607.03634 I was asked to contribute a chapter to Routledge Series on Philosophy of Rigor on the topic of AI. So here it is, my first paper in philosophy of science Hope I did alright. - My late PhD advisor used to say, "Too much rigor leads to rigor mortis " He was also the PhD advisor of John Jumper Nobel Prize 2024 and Moungi Bawendi Nobel Prize 2023 A good PhD program teaches not just a technical speciality but general scientific judgment. How Join the conversation