{"slug": "language-models-can-t-spark-scientific-revolutions-but-world-models-might", "title": "Language models can't spark scientific revolutions, but world models might", "summary": "Google DeepMind researcher Tom Zahavy argues in a position paper that large language models cannot spark scientific revolutions because they lack the cognitive mechanism of \"manipulative abduction\" — the ability to invent a cause for which no linguistic template exists. Zahavy contrasts this with Einstein's \"happiest thought\" about a freely falling observer, which came from embodied simulation rather than data-driven optimization.", "body_md": "# Language models can't spark scientific revolutions, but world models might\n\n**Can language models spark a scientific revolution? In a position paper titled \"LLMs can't jump,\" Google Deepmind's Tom Zahavy argues they can't. They're missing the cognitive mechanism needed to create something truly new.**\n\nZahavy builds his case on a framework Albert Einstein sketched in a letter to his friend Maurice Solovine. Discovery, Einstein wrote, is a cycle: sensory experience leads to an intuitive \"leap\" toward axioms, and from there, logical deduction produces testable conclusions. Axioms are the unproven foundational assumptions of a theory.\n\n## AI handles two of three types of reasoning\n\nTo pinpoint where the gap lies, Zahavy draws on a classic distinction from philosopher Charles Sanders Peirce, who categorized all reasoning by how it connects rules, cases, and results.\n\nDeduction derives guaranteed conclusions from fixed rules, like running a program that produces a provably correct output. Induction spots patterns in data: observe a thousand white swans, and you generalize that all swans are white. Abduction is the creative leap. It invents a cause to explain a surprising phenomenon.\n\nThis third form is where Zahavy sees the critical bottleneck, and he draws a line between two levels of it. Ordinary abduction picks the most plausible explanation from a set of known candidates, the way a doctor matches symptoms to a disease. Language models can do this, he concedes. The harder version is what he calls \"manipulative abduction\": inventing a cause for which no linguistic template exists yet. That, he argues, is the real bottleneck of scientific invention, and machines can't do it.\n\nInduction and deduction, the paper argues, are well within reach. Language models already excel at statistical pattern recognition, and they're rapidly conquering formal derivation too. Systems like AlphaProof, Gemini, and GPT-5 now achieve gold-level scores on International Mathematical Olympiad problems. Zahavy even concedes that a language model could probably derive general relativity if given Einstein's assumptions as a starting point. But formulating those assumptions in the first place, making the manipulative leap to reach them, remains the bottleneck.\n\nWhy machines struggle with this leap, Zahavy illustrates using that very theory: AI models typically learn by comparing their predictions to reality and adjusting based on the error, the gap between prediction and outcome. Without a detectable error, there's nothing for the system to work with. And that's the situation Einstein faced, Zahavy argues.\n\nWhen Einstein was working, there was no data crisis. Newton's physics had been confirmed with extreme precision. The only known anomaly, a tiny shift in Mercury's orbit, had been attributed to a hypothetical hidden planet called \"Vulcan.\" An optimization-driven AI would have had no reason to overthrow physics, Zahavy argues. Following the logic of the argument, it would have done what the astronomers of the era did: invented an extra planet to account for the small discrepancy, rather than rethinking space and time. The data confirming Einstein's theory, such as Eddington's measurement of light deflection, didn't arrive until years after the theory was formulated.\n\n## A jump requires a body\n\nSo where did the manipulative abduction come from that led Einstein to his axioms? Zahavy points to Einstein's \"happiest thought\": the freely falling observer who no longer feels gravity. This insight came from embodied simulation, Einstein mentally playing through a physical sensation rather than grinding through equations. He imagined a physicist inside an accelerating elevator in space and concluded that acceleration and gravity are indistinguishable from the inside.\n\nZahavy draws a parallel to Archimedes, who didn't discover his buoyancy principle through calculation but, as the story goes, through the physical feeling of water rising as he stepped into a bathtub. In both cases, a foundational principle emerged that didn't yet exist in the language of the time.\n\nLanguage models lack exactly this sensory grounding. Zahavy compares them to philosopher John Searle's \"Chinese Room,\" a famous thought experiment where a person shuffles Chinese characters according to a rulebook without understanding a single word. Language models shuffle the symbols of physics in much the same way, without access to the physical experience that gives those symbols meaning.\n\nSakana's AI Scientist and Deepmind's AlphaEvolve automate scientific workflows impressively. But the AI Scientist only recombines existing concepts, while AlphaEvolve optimizes brilliantly yet needs a clear error signal it can shrink step by step. Einstein never had that signal. Neither system, Zahavy argues, can make the leap into an entirely new framework of thought.\n\n## World models as a path to abduction\n\nAs a possible way forward, Zahavy points to physically consistent world models. He draws a line here: video generators like [Veo](https://the-decoder.com/deepmind-says-video-models-for-visual-tasks-could-become-what-llms-are-for-text-tasks/) simply predict the most likely next frame. A falling apple falls not because the model understands gravity, but because falling is the most common continuation in the training data. That's still just pattern matching.\n\nAction-controllable world models like [Genie](https://the-decoder.com/google-deepmind-opens-project-genie-to-us-subscribers-for-real-time-ai-world-generation/), on the other hand, let an agent actively intervene in a simulation and run counterfactual experiments, like mentally cutting an elevator cable. A \"synthetic lab\" like this could provide the feedback loop needed to invent new axioms where no linguistic template exists yet.\n\n```\nAI News Without the Hype – Curated by Humans\n\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive \"AI Radar\" frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t\n\n\t\t\t\t\tSubscribe now\nRead on for the full picture.Subscribe for hype-free coverage.\n\nAccess to all THE DECODER articles.\nRead without distractions – no Google ads.\nAccess to comments and community discussions.\nWeekly AI newsletter.\n6 times a year: “AI Radar” – deep dives on key AI topics.\nUp to 25 % off on KI Pro online events.\nAccess to our full ten-year archive.\nGet the latest AI news from The Decoder.\n\nSubscribe to The Decoder\n```\n\n", "url": "https://wpnews.pro/news/language-models-can-t-spark-scientific-revolutions-but-world-models-might", "canonical_source": "https://the-decoder.com/language-models-cant-spark-scientific-revolutions-but-world-models-might/", "published_at": "2026-07-30 14:01:33+00:00", "updated_at": "2026-07-30 14:11:37.768540+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["Google DeepMind", "Tom Zahavy", "Albert Einstein", "Charles Sanders Peirce", "AlphaProof", "Gemini", "GPT-5"], "alternates": {"html": "https://wpnews.pro/news/language-models-can-t-spark-scientific-revolutions-but-world-models-might", "markdown": "https://wpnews.pro/news/language-models-can-t-spark-scientific-revolutions-but-world-models-might.md", "text": "https://wpnews.pro/news/language-models-can-t-spark-scientific-revolutions-but-world-models-might.txt", "jsonld": "https://wpnews.pro/news/language-models-can-t-spark-scientific-revolutions-but-world-models-might.jsonld"}}