cd /news/ai-research/on-the-existence-of-non-sofic-groups · home topics ai-research article
[ARTICLE · art-127068] src=terrytao.wordpress.com ↗ pub= topic=ai-research verified=true sentiment=↓ negative

On the existence of non-sofic groups

OpenAI's paper claiming ten advances in mathematics, including a proof of the existence of a non-sofic group, relied on a 2019 paper by Andreas Thom and Gábor Kun in its crucial Proposition 2.3, Thom wrote, after the company's public announcement initially framed the field as having made "no progress" in the last decade. Thom said he emailed Mark Sellke and Sébastien Bubeck calling the framing "intellectually dishonest," and Sellke agreed to a revision that changed the announcement. Thom also asked whether his months of discussions with ChatGPT about expander decompositions and centralizer rigidity were part of the model's training data or accessible to its reasoning process, citing an "unacceptable" lack of transparency.

by read10 min views1 publishedSep 11, 2026
On the existence of non-sofic groups
Image: Terrytao (auto-discovered)

[This is a guest post by Andreas Thom. This blog post was initially written in a different file format and converted using AI. — T.]

When I woke up on August 1st, 2026, I had received a few emails from colleagues asking for my opinion on a remarkable result that had circulated the previous day. The result was a solution to a long-standing open problem in geometric group theory, specifically the existence of a non-sofic group. I was astonished and at the same time, looking at the first draft, also in a way happy to see that Kun’s work on expander decompositions and my joint work with Gábor Kun played a decisive role in the crucial Proposition 2.3 of the OpenAI paper. I had always hoped that the theory of centralizer rigidity would eventually have significant applications, but I had not found the right setting in which it could be used so effectively. In that sense, the solution also came as a relief and I was happy to explain the ideas in a post on MathOverflow a few days later.

From the start, colleagues pointed out that the framing in the public announcement that appeared shortly afterwards was misleading, in that it spoke of “no progress” in the last decade, while relying on our 2019 paper (not to mention subsequent work by many hands that was not directly relevant for the OpenAI paper but would still be considered to be progress by many). So I wrote to Mark Sellke and Sébastien Bubeck: “[…] I find the framing intellectually dishonest. You (and I am talking about you personally, since I have no one else to address this to) cannot speak in the public announcement of a decade without progress and then use a 2019 paper in a crucial way. It is true that Proposition 2.3 is a really clever use of the centralizer-rigidity theorem, but neither does its short proof require new techniques […].

“It is true that the last stone finishes the building and usually those who can put it get the credit for solving the problem, that is fair enough. I have no problem with that and I personally do not care much about credit. However, I guess you would get enough praise without downplaying the previous contributions.”

Sellke replied to this and basically agreed to the need for a revision; as a result the public announcement was changed to the form it has now. I was glad to have received an early draft from Sellke also on August 1st, otherwise there would have been no way to react to the first public announcement at all, since neither the PDF nor the website of the announcement contained contact information. Anyway, I was happy that this was resolved and the matter closed.

It is fair to say that the approach of Kun and myself had not been viewed as the main line of attack on non-soficity prior to OpenAI’s announcement. In fact there were other more promising approaches along the line of quantum games etc. at the time, that had already led to a negative solution of the famous Connes Embedding Problem and, later, the disproof of the Aldous–Lyons conjecture. Hence, OpenAI’s detailed command of the techniques of Kun and myself made me wonder how the model found this route, especially since I discussed these techniques and their use in extensive sessions with ChatGPT over the last months.

So in the same email I asked Mark Sellke and Sébastien Bubeck: “Another point is that I and a colleague in Dresden were discussing the expander matching problem and various extensions of the work with Gábor Kun actively over the last months with ChatGPT, so that we are of course curious if that was part of the training data or accessible to the reasoning process. There is a certain (frankly unacceptable) lack of transparency here; and I fear it will damage the communal process of math more than the new AI-generated results will benefit the subject.”

Mark Sellke’s complete answer to this part of my email was: “Regarding your conversations with ChatGPT: that did not happen.”

Anyway, I thought, these techniques were public, so their use is not evidence that our conversations influenced the model. But because this was not the main line of attack, and because I had recently discussed precisely these techniques and possible extensions with ChatGPT at length, I thought the question had to be asked. Back at the beginning of August, I then returned to mathematics and wrote a subsequent paper with Gábor Kun on applications of the ideas that were the basis of Proposition 2.3. This was my way to react; after all, the integration of the new result in the math landscape seemed like a natural next step.

However, after reading up on the controversy around the Buckmaster–Alpöge case, the whole story came back to me and I realized that the answer I received from OpenAI was misleading, to say the least. I already wrote about this briefly on Mathstodon.

I had explicitly asked about two different things: (1) whether our conversations entered training data, and (2) whether they were accessible to the solving process. OpenAI said in the Buckmaster–Alpöge case that no specific user data was accessed, but added that it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” In light of OpenAI’s later wording, I cannot tell whether Sellke’s answer denied both possibilities or only direct access under (2). No qualification, explanation, or evidence was given. Whatever its intent, I regard the answer as materially misleading.

OpenAI was drawing a distinction that its answer to me erased, despite the fact that my question explicitly made that distinction. We are not required to reverse-engineer OpenAI’s internal training pipeline to establish what happened. Only OpenAI has the relevant data for that. For such a categorical denial by OpenAI to be credible, OpenAI should disclose its basis: product and privacy settings, relevant datasets and checkpoints, and what “de-identified data derived from usage” means.

I disabled model training on 29 June. That control is still only a promise whose implementation users cannot audit, and it is prospective: it does not answer what happened to earlier conversations or to derivatives already selected.

If nonpublic research supplied by users improved a model and the provider then used that model to race those users to publication—without informed consent, disclosure, or credit—that would be ethically indefensible. De-identification may remove a name; it does not remove the intellectual content of a mathematical idea. Sellke and Bubeck seem to be blind to this simple moral aspect. Sellke gave me a categorical assurance without explaining its basis; I regard that response as materially misleading. If he lacked the information needed to rule out training use, he had no basis for giving that assurance. Bubeck’s acknowledged career-related remark in the conversation with Buckmaster and his objection to including Alpöge in a proposed paper presenting OpenAI’s proof deepen my concern about their commitment to academic standards. Taken together, these episodes raise serious questions about their judgment and personal integrity.

I am not claiming that anyone read individual chats or that our conversations were in fact used in training; I do not know that. My criticism concerns the categorical denial. If they did not know what entered the training data (the most likely scenario), they should have said so.

After I finished writing this post, I received a message from Mark Sellke, who acknowledged understanding how I “reasonably arrived at [my] conclusions given the evidence available”. He pointed me to a discussion citing OpenAI’s new statement that Buckmaster’s Codex prompts from the preceding two months could not have influenced its system, including through training. I wish I could trust this more. In any case, it suggests that the math community can successfully put pressure on the industry to take these issues at least somewhat more seriously.

So what does that all mean and how do we as a community proceed? Setting aside these particular cases (which might also be very different in what really happened behind the scenes), we have to see the broader picture and I believe there is no way of going back.

As far as I see it, a mathematical publication used to bring together three things. It announced a result, identified the people who had produced it, and added something to our human understanding of mathematics. It could therefore serve at the same time as a record of knowledge, a basis for assigning credit, and evidence of a mathematician’s ability. AI breaks this connection.

An AI system can produce a correct proof even if no human being discovered or even understood the argument in the usual sense. A typical form that this can take nowadays varies from somewhat verified AI Slop to a Lean certificate, or a combination of both. The proof may be worth publishing, but the publication then records only that the result has been established. It does not necessarily tell us how the result was found or who understood it.

The question of credit and contribution may have no satisfactory answer. A model may draw on published work, feedback, conversations, prompts, and its own search in ways that apparently cannot be reconstructed clearly. The person or the company who ran the model should not simply receive whatever credit cannot be assigned elsewhere. As another consequence, a publication record can no longer serve as a reliable measure of a person’s mathematical quality. If the theorem, proof, and written explanation may all have been generated by AI, a list of papers tells us very little about what the named author contributed or understands. Even if matters of data privacy are resolved, I see only little hope that the current system of publication and credit can be salvaged. The whole idea of personal credit will not work in such a highly connected environment anymore and I actually think that the focus on “who got something first” (not to speak of prizes for a solution to particular problems) was always misleading, even though a powerful driving force.

The more important question is therefore what someone contributes to human mathematical understanding. This includes explaining why an argument works, separating the main idea from technical details, connecting a result with other areas, finding the right questions, teaching new methods, and helping other mathematicians make use of them. Such contributions may happen through papers, but also through lectures, discussions, teaching, and collaboration.

A formal proof certificate is comparable, in a sense, to the detection of a new star. It confirms that something is there, but further work is needed to understand what it is, why it matters, and where it belongs in the larger landscape. Mathematics is only to a lesser extent the production of correct statements. It is more the human process of making sense of them.

The institutional issue goes far beyond mathematics. Advanced AI is becoming a general supply of “intelligence” on which science, education, public administration, industry, and ordinary life may all depend. It should therefore be provided according to standards comparable to those governing water or electricity: reliable and broadly available, with clear public duties, strong privacy rules, independent oversight, and protection against discriminatory or self-serving use.

Intelligence of this kind should not be treated as an ordinary consumer product whose conditions are set entirely by a few companies. A provider should not be able to collect people’s ideas and information, control all evidence about how they were used, and then exploit that advantage against its own users. Once intelligence becomes basic infrastructure for society, it must be highly regulated and governed in the public interest.

── more in #ai-research 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/on-the-existence-of-…] indexed:0 read:10min 2026-09-11 ·