Yesterday I sat down with GPT 5.6 Sol High to do some brainstorming. The topic was one of the less appreciated Millenium Problems (the Birch and Swinnerton-Dyer conjecture), and the initial prompt started life as a question on Quora.
Number theory is not a field I know, and I only expected the "discussion" to last for a few exchanges. However, ChatGPT was immediately inspired to try generalizing the BSD conjecture in a specific direction, and this led to an unusually drawn-out and self-sufficient line of "research". Almost every response concluded with a suggestion as to what the next task should be, and my input was just to cheer on what had been accomplished so far, and then endorse the suggested next direction.
What was especially striking to me, was the frequency with which each new response began with a conceptual adjustment regarding the sub-task to be performed. Evidently a vast variety of abstract objects are possible in number theory, and their differences and interrelations can be quite subtle. ChatGPT was regularly adjusting the next sub-task it had set itself, generally in the direction of greater nuance by aiming at a more sophisticated construction than it had first planned.
I was not able to judge what was going on with an expert eye, but it was a kind of interaction I had not quite had before. I have had lengthy brainstorming chat sessions with AI before, but mostly in physics, an area where I know something and could participate as an equal. Here, the AI's line of thought was all but self-sufficient - as I have mentioned, my role was largely just to say "yes, take the next step you just suggested" (I hear that vibe coding can be like this) - and, there seemed to be a lot of creative improvisation along the way, it wasn't just following well-established technical procedures of the field. (But this is where I could be wrong - perhaps it's normal, when problem-solving in advanced number theory, to have to make numerous nuanced judgment calls regarding the specific type of object that you are trying to construct.)
As the end approached today (the end being, arriving at the new generalized form of the conjecture), I felt certain that I was witnessing something that was worth posting about... Then on X, I noticed that today's news included an announcement by Anthropic that one of their employees had last week managed to radically improve a bound related to the Riemann hypothesis, and did so by urging Claude to be ambitious and believe in itself. Oh well. Perhaps I should be satisfied that for $30/month, I got to experience an echo of what the frontier labs get to do, with their billions of dollars and legions of PhDs.
My real message is that I believe most people are underestimating the significance of the recent AI successes in research-level math. This is some of the hardest thinking that humans can do, and it is now being automated. That is not a capability that will remain bottled up in the realm of pure math, affecting only mathematicians. I am much more inclined to think that this is one of the very last signs before AI becomes smarter than humans at absolutely everything.