Why math wants the whole chain open The Advisory Group on Mathematics and Artificial Intelligence (AGMAI) published guidelines asking frontier AI labs to stop testing advanced mathematical problems on proprietary models, stating "we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." The objection follows OpenAI's release of mathematical results produced by "an unreleased internal OpenAI model," which the group says leaves no second system to check against, unlike the 2014 case where Mathematica and Maple disagreed on determinant computations and Maple and Sage agreed on the correct answer. The guidelines argue mathematics sits at the farthest end of the open-versus-closed spectrum because the end-to-end chain, including how an answer was reached, must be open. In The flood, released responsibly? ../math/openai-release.html I quoted the opening of the guidelines https://agmai.org/general-sep29/ that the Advisory Group on Mathematics and Artificial Intelligence AGMAI published for AI labs: At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. … we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models. OpenAI’s release came from “an unreleased internal OpenAI model”, and it intends to keep evaluating its internal models on mathematics. But science uses proprietary tools all the time. So is the objection that the models are proprietary, rather than that frontier AI is used on advanced math at all? And why should math be held to a stricter standard than the rest of science? I think one way to see it is that math, on the spectrum of open vs. closed “dependencies”, is probably on the farthest end: absolute openness. E.g. that’s why you can’t copyright https://www.law.cornell.edu/uscode/text/17/102 math, as far as I understand. Even in my field, while our science is open, the software used a dependency of our science might be proprietary. E.g. industrial, proprietary, commercial software simulating electromagnetic effects is used, and is core, to design the amazing sinuous antennas https://arxiv.org/html/1210.8256v1/AntCenter.png of the multichroic detectors in CMB experiments. The ones in Suzuki et al. 2012 were designed with HFSS, a commercial 3D electromagnetic simulator. So part of the design process is not entirely open: the theory behind it is, the software used in the design is not, and the effect of the design is measured in the lab, so it is open again. This has always been a bit unsettling to me, perhaps to my inner mathematician. But it seems to be normal in science: we use SOTA technology to advance science, and as long as the “scientific” part is open, we tolerate it. In math, the end-to-end chain “must” be open, and that should include how you came up with the answer as well. So an open model is best, but even a proprietary model should at least be publicly available. Using an internal model means only someone from OpenAI can do that, which is the farthest from that ideal, kind of the opposite of the democratization of math. This is not entirely new. Durán et al. 2014 were using Mathematica to look for counterexamples to their conjectures generalizing a result of Karlin and Szegő on orthogonal polynomials, which came down to computing determinants of matrices of big integers. Mathematica found some counterexamples: Fortunately, another of us was using Maple, and when checking those supposed counterexamples he found that they were not counterexamples at all. Mathematica was computing the determinants wrongly, and gave a different answer when the same determinant was evaluated twice. Maple and Sage agreed on the right one. And their conclusion is the same unsettling feeling: The commercial computer algebra systems are black boxes and their algorithms are opaque to the users of course, also the source code , and certainly this does not contribute to avoid errors. … Moreover, known bugs of computer algebra systems should be available to the users; this is usual in free software, but an anathema for commercial packages. Using more than one system is a way to reduce the trust in a single proprietary software, and interestingly, comparisons like this sometimes uncover hidden bugs. But that only reduces it. Again, the scale and reach of LLMs make that issue more prominent and intolerable, and with an internal model there isn’t even a second system to check against. AGI, where the G is general, cuts both ways: benefits and harms are both generally available and you’re welcome .