cd /news/artificial-intelligence/mathematicians-may-be-worried-but-ai… · home topics artificial-intelligence article
[ARTICLE · art-83317] src=lesswrong.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Mathematicians may be worried, but AI-for-science is going to be great, recursively self-improving, and we’re going to learn loads

Simon DeDeo, a researcher at Carnegie Mellon University, reports that AI models such as GPT 5.6-sol can solve open mathematical problems with at least a 1% success rate when prompted with a problem, and he has used the model to extend three of his own papers, finding the results promising for AI-for-science. DeDeo's experiments, conducted as part of the Templeton-funded Proofs & Reasons project, involved giving the AI unrestricted access to computational resources and letting it run for a day or two, leading to new insights and hypotheses.

read7 min views1 publishedAug 1, 2026

I have a new post on my experiments with 5.6-sol as a scientific agent; you can find the original post here. I've reproduced it below, as well. I'd be grateful for any feedback, thoughts, or to hear about experiments other people have done; you're very welcome to point your own agents at the repos described below, or anywhere else you like.

As part of our Templeton-funded Proofs & Reasons project, I travelled to the ICM in Philadelphia this year, to help a collaborator run some new experiments on expert mathematicians. The ICM is an every-four-years event that you may remember as where Hilbert launched his field-defining 23 problems, or the place they announce the Fields Medals. It’s a big deal for mathematics, and we had a wonderful time both running our experiments and talking to the truly international community whose interests and abilities define modern mathematics.

At the same time that ICM was running, it was hard to ignore the vast number of AI-enabled results that were breaking, almost simultaneously, on Twitter and elsewhere. Commercially-available models were cracking, in hours, problems that had stood for decades. There are too many to summarize, but this Twitter pairinggives you a sense of both the advances AI has made, and the responses from some human mathematicians. At this point, “google for an open problem and stick it into GPT 5.6-sol on high” seems to have at least a 1% success rate—which fills some mathematicians with hope (new proofs from the alien proof machine!) and others with a bit of despair (new proofs… from an alien proof machine!)

It’s natural to ask what this looks like for science. We don’t have well-framed open questions in the same way mathematicians do, but we do have plenty of good problems. I decided to sic GPT 5-6 sol (on high) on three of my own papers, with a generic instruction to do something cool. I picked three papers that (1) I’m particularly proud of, and (2) that seemed amenable to this kind of study—empirical data science investigations with a strong theoretical backing.

These papers were:

My initial prompts said, roughly, “read the paper, and extend it in an interesting direction, here are some resources, please ground your final report in the literature.” Here’s the initial prompt for Murdoch et al.:

I’d like you to read http://XXX (also a PDF in this folder, and a separate PDF with supplementary). This is a project that uses some simple methods to study how a creative intellectual explores and searches the space of his intellectual world on the way to producing a major intellectual revolution.

I’d like you to extend and this analysis to other intellectuals, artists, writers, philosophers, and scientists to see how well these generalize. You can also search the cog sci, biographical, and other literatures for useful for new hypotheses and questions.

To aid your work, you have passwordless ssh access to XXX.lan.cmu.edu, including passwordless sudo, a high powered 32 core machine with lots already installed; you also have access to the CMU ORCHARD cluster, which has additional resources, including storage.

Work carefully, use all the resources you can, and seek out both new data and new hypotheses. We’re looking for common themes that might help us understand how great minds work, potentially with deeper lessons for more “ordinary” people. Work assiduously and carefully, but also creatively and with an eye towards fundamental and basic questions.

I let all three of them run, simultaneously, in three VS Code windows over the course of a day or two. I disabled all safety mechanisms, and gave it unrestricted access. CMU IT services blocked my IP because of suspicious traffic — mostly lots of traffic to my local cluster and ORCHARD — but I pled my case and was unblocked.

I intervened occasionally, asking questions, making suggestions, and occasionally giving orders. The three projects consumed millions of tokens, and thousands of CPU hours on machines running different pieces of analysis code in parallel. GPT was instructed to produce a research report with a <250 word abstract and a max of 10 pages (modulo bibliography and appendices), and then to audit and correct it repeatedly for clarity and validity.

Here were, roughly, the three outcomes:

All of this is incredibly exciting. I have no idea if these results are real — please don’t take them as such! — and the papers leave a lot to be desired, but the sense checks suggest that none of them are completely crazy. Both GPT-Murdoch-etal and GPT-Viteri-etal, for example, replicated our original results (and pointed out how they might have made different choices, thanks clankers), and GPT-Viteri-etal found subtleties in going from Coq to Lean that showed it was clearly able to make sense of a complex story.

In all three cases, 5.6-sol located an interesting and promising direction for investigation that would, if fleshed out and developed, be not only a publication-worthy piece of work (i.e., novel enough, and well-sourced enough to pass the bar at a good journal), but something that legitimately increased our species’ store of knowledge: about how we interact with our cultures, how we govern ourselves democratically, or how AI is building mathematical knowledge in a new (and potentially suboptimal) fashion.

I honestly don’t see a downside here, at least as things currently stand. The machines are producing intriguing discoveries that are simple and clear enough to be the target of human investigation, and doing so in a way that is pleasurable and fun to monitor and interact with. The artifacts it produces are auditable and clear — all on github — and it was aggressive (but not insane) in pulling in data from a wide variety of sources. It’s on me (or anyone else) to ask if anything in (1), (2) and (3), above is actually worth pushing further on.

There’s also a clear suggestion that we’re entering an era of recursive self-improvement. For example, GPT-Barron-etal relied upon https://github.com/simon-dedeo/temporal-lda — a project I created with Claude a few months ago that is a modified version of a classic algorithm, rewritten (and parallelized!) to handle uneven temporal coverage, a classic problem when running Bayesian inference on historical data that nobody had taken the time to do. The new githubs from this little experiment are now also publicly available, so all those CPU hours, and my obviously genius executive decisions, have entered the creative commons for the next machines to rely upon. All those proto-discoveries, and tools, are available to the next machines that tackle projects — potentially very different ones from those that my collaborators and I were interested in. There are lots of things to manage here, like data poisoning, which Bálint Gyevnár is working on, but while these are critical problems, none of them seem insurmountable (so please keep funding those people).

What I’m finding matches what I’ve heard from my old advisor (and now Simons Foundation leader) David Spergel: we are all — whether we’re graduate students, or senior faculty — managing a private research group now. Graduate training will likely have to pivot to handle this new challenge, and evaluation of more early-career colleagues will need to shift from “can you do task X?” to “what questions did you get agent Y to answer, and how well did you validate them?” We’ll be less impressed by technical tours de force, and more impressed by depth, meaning, and (why not!) pleasure. What’s not to like?

One critical research question is what happens when you repeat this process a hundred times; that’s above the scale that I can monitor and guide, but that’s just another layer in the org chart from the Spergel persepctive. The results above are a one-time short; GPT-Murdoch-etal was probably the most impressive, but perhaps GPT-Barron-etal and GPT-Viteri-etal just needed a fresh start (go for a walk, robot!) There’s a huge amount of hysteresis and once a clanker has a hold of an approach, it’s hard to dislodge them. GPT-Murdoch-etal, for example, had some lovely preliminary results on dwell-time, but just didn’t want to rewrite the abstract.

The longer timescale horizon is more intriguing yet. In ten years, will be we able to sort the wheat from the chaff in the flood of githubs that will likely emerge from this kind of work? How will we direct our scientific attention? How firm will our knowledge be? Here we’re in a worse position than the mathematicians, I think — they, at least, live in a world of deduction, and they have tools like Lean. But what will we scientists use to validate their inductive and abductive knowledge?

PS sorry that Claude Mythos stole your credit card.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @simon dedeo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mathematicians-may-b…] indexed:0 read:7min 2026-08-01 ·