{"slug": "fields-medallists-tell-ai-labs-to-stop-testing-maths-on-models-nobody-else-can", "title": "Fields medallists tell AI labs to stop testing maths on models nobody else can use", "summary": "The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), an independent nine-member group linked to the Institute for Advanced Study in Princeton whose members include Fields medallists Timothy Gowers, Martin Hairer and Edward Witten, published guidelines on Tuesday asking AI labs to stop testing advanced mathematical problems on proprietary models and to stop treating maths results as marketing vehicles. The guidelines, built on more than 600 survey responses and one case in which OpenAI \"announced the existence of many results without giving details,\" require labs releasing results no human understands to cite prior literature, publish the model name, prompts, summarised chain of thought, time taken and estimated compute cost, formalise proofs where possible, and report how many similar problems the model failed. AGMAI also says labs must fund conferences, summer schools, postdocs and expository books through existing non-profit institutions, and warns that internal frontier models risk \"a two-tier system where labs outrun the rest of the field.", "body_md": "Timothy Gowers speaking at the GOSIM conference at Station F in Paris in May 2026, under a slide titled “What it feels like to be a mathematician today”. Image: [Bretwa](https://commons.wikimedia.org/wiki/File:Timothy_Gowers_at_GOSIM_conference_at_Station_F_in_Paris_-_2026-05-05.jpg) / Wikimedia Commons, CC0, cropped\n\nSome of the world’s most celebrated mathematicians have told AI labs to stop testing hard maths problems on private models nobody else can use. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), whose members include Fields medallists Timothy Gowers, Martin Hairer and Edward Witten, [published guidelines on Tuesday](https://agmai.org/general-sep29/) for how labs should release AI-generated maths, and opened with a blunt request: “we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.”\n\n## Who is behind it\n\nAGMAI is linked to the Institute for Advanced Study in Princeton. According to [its own website](https://agmai.org/), it began when OpenAI approached some of its members about setting up an external advisory board, and they chose to form an independent group instead. Its nine members are François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten and Melanie Matchett Wood.\n\nThe group says it asked mathematicians what responsible release should look like and got more than 600 replies. The survey was built around one real case: a time when OpenAI “announced the existence of many results without giving details.” The recommendations apply to any lab whose models are likely to have a significant impact on mathematics, not just OpenAI.\n\n## “Human understanding of mathematics remains of paramount importance”\n\nThe core worry is that AI can now produce proofs that the person who prompted it can’t understand, check or take responsibility for, which breaks one of the oldest rules in mathematics. So the group splits AI results in two. If a mathematician fully understands a paper, it should go through the usual route: a preprint, peer review and talks. If nobody understands it yet, the lab has extra work to do:\n\n- **Credit the humans:** search the literature and cite where the ideas first appeared, even if the model found them on its own.\n- **Write it properly:** produce a readable version, not something “full of wordy reasoning and non-standard terminology.”\n- **Show the working:** publish the model’s name, the prompts, a summarised chain of thought, the time taken and the estimated compute cost.\n- **Formalise it:** where possible, check the proof in a formal system, or say clearly that it hasn’t been.\n- **Admit the misses:** when releasing a batch of results, say how many problems of similar difficulty the model tried and failed on, and how they were picked.\n\nResults should also go into scholarly repositories that no AI lab controls. The group “strongly” recommends that labs stop treating maths results “as marketing vehicles to promote their models.”\n\n## Labs should pay for the catch-up\n\nThe guidelines say labs that release maths nobody understands “must take responsibility for ensuring that human understanding will follow,” including with funding for conferences, summer schools, postdocs and expository books. But the labs shouldn’t pick what gets funded: that should be left to existing non-profit institutions. And the group is clear that taking the money “would not be conferring legitimacy on the practices of the AI labs.”\n\nIt also warns that frontier results produced on internal models risk “a two-tier system where labs outrun the rest of the field,” and asks labs to give mathematicians everywhere broad, equal access to their public models.\n\nThe guidelines land after weeks of friction between labs and mathematicians, including [the row over OpenAI’s Navier-Stokes proof](https://madrobot.blog/2026/09/19/mathematicians-threatened-by-ai-cant-quit-it/). OpenAI’s [latest models](https://madrobot.blog/2026/09/29/openai-devday-2026-everything-announced-dots-gpt-6-1-sol-pro-500/) and the [cancelled GPT-6.1 Astra](https://madrobot.blog/2026/09/28/openai-scraps-gpt-6-1-astra-release-safety-concerns/) were all pitched partly on their maths ability. OpenAI hasn’t publicly responded to the recommendations.\n\n## Why it matters\n\nMaths is one of the few fields where AI labs can point to results that are provably right, which makes it their favourite showcase. When the people who set the field’s standards say the showcase is being run the wrong way, labs that want mathematicians’ trust, and their help checking the results, will find it hard to ignore.\n\n*Sources: [AGMAI, “Responsible Release of AI-Generated Mathematics”](https://agmai.org/general-sep29/) ([PDF](https://agmai.org/wp-content/uploads/2026/09/recommendations.pdf)), [AGMAI](https://agmai.org/)*", "url": "https://wpnews.pro/news/fields-medallists-tell-ai-labs-to-stop-testing-maths-on-models-nobody-else-can", "canonical_source": "https://madrobot.blog/2026/09/30/fields-medallists-ai-labs-stop-testing-maths-proprietary-models-agmai/", "published_at": "2026-09-30 07:02:00+00:00", "updated_at": "2026-09-30 07:19:06.593530+00:00", "lang": "en", "topics": ["ai-policy", "ai-safety", "artificial-intelligence", "ai-research"], "entities": ["Advisory Group on Mathematics and Artificial Intelligence", "Timothy Gowers", "Martin Hairer", "Edward Witten", "Institute for Advanced Study", "OpenAI", "François Charles", "Melanie Matchett Wood"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fields-medallists-tell-ai-labs-to-stop-testing-maths-on-models-nobody-else-can", "markdown": "https://wpnews.pro/news/fields-medallists-tell-ai-labs-to-stop-testing-maths-on-models-nobody-else-can.md", "text": "https://wpnews.pro/news/fields-medallists-tell-ai-labs-to-stop-testing-maths-on-models-nobody-else-can.txt", "jsonld": "https://wpnews.pro/news/fields-medallists-tell-ai-labs-to-stop-testing-maths-on-models-nobody-else-can.jsonld"}}