{"slug": "the-role-of-rigor-in-artificial-intelligence", "title": "The Role of Rigor in Artificial Intelligence", "summary": "A new paper by an AI researcher argues that the common criticism of deep learning as 'alchemy' misses the mark, proposing instead a framework of three forms of rigor—conceptual, epistemic, and operational—to explain AI's progress and uncertainty. The paper highlights that while operational rigor (benchmarks, evaluations, deployment controls) is strong, conceptual and epistemic rigor lag, leaving basic questions about understanding, generalization, and failure unanswered. The author calls for balanced development of all three forms to mature AI as both a science and a technology.", "body_md": "“Deep learning is alchemy” may be the most repeated criticism in AI. It also misses the mark.\nAlchemy failed to deliver results. Deep learning, by contrast, has produced transformative technologies. And fields like medicine are only partially understood without being deemed alchemical.\nSo calling AI “alchemy” captures part of the problem, but not all of it. Modern AI is not simply undisciplined experimentation. It contains significant amounts of rigor. But we still struggle to answer basic questions:\n• Do models understand?\n• Why do they generalize?\n• When will they fail?\nThe deeper issue is that rigor takes different forms—and in AI, those forms are unevenly developed.\nMy new paper distinguishes three:\n• Conceptual rigor: coherent terminology and paradigms\n• Epistemic rigor: reliable scientific understanding\n• Operational rigor: reliable performance and deployment\nThis framework helps explain both the extraordinary progress of modern AI and the uncertainty surrounding it.\nConceptual rigor asks whether the field knows what it's talking about.\n• What exactly is intelligence?\n• What qualifies as AGI?\n• What does it mean for a system to be aligned?\nConsider the debate over whether current models are intelligent. One person points to their breadth of performance. Another points to weak planning. Another emphasizes sample inefficiency. Another asks whether it has a grounded model of the world.\nThey appear to disagree about one property. Often, they are evaluating four.\nThis is why conceptual clarity matters in practice. Questions about intelligence, understanding, AGI, and alignment do not remain confined to philosophy: they shape how things are measured, optimized, and built.\nEpistemic rigor asks whether empirical success has become scientific understanding.\nThe paper focuses on three criteria:\n• Can findings be reproduced?\n• Can behavior be predicted in advance?\n• Can success and failure be explained?\nAI experiments are unusually reproducible in principle: code, data, and models can be copied. But conclusions may still depend heavily on random seeds, hyperparameters, implementation choices, benchmark selection, and compute budgets.\nReproducing a number is not always the same as reproducing the conclusion drawn from it.\nPrediction is harder.\nScaling laws can forecast some training outcomes. Infinite-width theory can lead to more tractable settings. Classical learning theory explains important pieces. But we still lack broad principles telling us when a model will generalize, fail under distribution shift, or remain robust under adversarial perturbations.\nExplanation is harder still.\nNeural networks are mathematically specified, yet their learned features resist human interpretation. A behavior may arise from training data, optimization dynamics, internal representations, or interactions among all of them. The system is transparent in code but opaque in meaning.\nOperational rigor is where modern AI is strongest: benchmarks, evaluations, monitoring, red-teaming, and deployment controls.\nThe field has become highly effective at improving systems without first obtaining a scientific theory of them. Benchmarks turn capabilities into measurable targets. Post-training shapes behavior. Tools and scaffolding compensate for model weaknesses.\nOperational rigor can therefore partially substitute for scientific understanding. That imbalance defines the deep-learning era:\n• Capabilities rise rapidly.\n• Explanations lag behind.\n• Benchmarks become optimization targets.\n• New systems generate new phenomena faster than theory can absorb them.\nAI is advancing while continually changing the object that science must explain.\nFor AI to mature as both a science and a technology, it will require all three forms of rigor:\n• Clearer concepts to define our goals.\n• Stronger science to predict and explain system behavior.\n• Better engineering to make systems genuinely reliable.\nThe future of AI depends not simply on demanding “more rigor,” but on identifying which kind is missing—and understanding how the imbalance shapes what we can build, know, and control.\n\n- Paper link:\n[arxiv.org/abs/2607.03634](https://arxiv.org/abs/2607.03634)I was asked to contribute a chapter to Routledge Series on Philosophy of Rigor on the topic of AI. So here it is, my first paper in philosophy (of science)! Hope I did alright. - My late PhD advisor used to say, \"Too much rigor leads to rigor mortis \" He was also the PhD advisor of John Jumper (Nobel Prize 2024) and Moungi Bawendi (Nobel Prize 2023) A good PhD program teaches not just a technical speciality but general scientific judgment. How\n# Join the conversation", "url": "https://wpnews.pro/news/the-role-of-rigor-in-artificial-intelligence", "canonical_source": "https://twitter.com/iamtimnguyen/status/2077091364743807065", "published_at": "2026-07-31 08:48:54+00:00", "updated_at": "2026-07-31 09:22:52.966058+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-safety"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/the-role-of-rigor-in-artificial-intelligence", "markdown": "https://wpnews.pro/news/the-role-of-rigor-in-artificial-intelligence.md", "text": "https://wpnews.pro/news/the-role-of-rigor-in-artificial-intelligence.txt", "jsonld": "https://wpnews.pro/news/the-role-of-rigor-in-artificial-intelligence.jsonld"}}