Examples of problems that current AI models fail to solve A MathOverflow user is seeking examples of mathematical problems that current AI models, including the best LLMs from OpenAI and Anthropic, fail to solve in a reasonable time or solve incorrectly without human assistance, yet average specialist human mathematicians can solve. The user notes that positive AI problem-solving results are overreported due to observation bias, and that problems previously solved by humans may already be in training sets, making it difficult to assess true AI capabilities. The question aims to establish upper bounds on AI mathematical abilities beyond trivial failures like counting 'r's in 'strawberry'. Background and motivation: There is now a long list of problems https://mathoverflow.net/questions/502120/examples-for-the-use-of-ai-and-especially-llms-in-major-mathematical-development , including many long-standing conjectures, that have been solved autonomously with current AI models. The problem with this sort of list is that it conveys the idea that LLMs are now far better than humans at problem-solving because they are able to solve many problems that humans couldn't: while the conclusion might be correct, the reasoning certainly is not: it might also be the case that problem solving difficulty for LLMs and humans is not well-correlated, and selecting problems hard for humans still leaves a lot of problems that are easy for AI selection bias . Underlying this is also the huge observation bias that “AI model solves long-standing conjecture” is considered newsworthy and attentionworthy whereas “AI model fails to solve problem that turns out to be not that difficult after all” won't get reported. There is the additional difficulty that any problem previously solved by humans is probably already in the AI models' training set. I am aware of the First Proof project https://1stproof.org/ which seeks to confront this observation bias in a more scientific way, but I think it is still interesting to gather anecdotal evidence to put the positive AI news into perspective. Anyway, if we are to navigate this brave new world that has such AI in't, we need to understand its abilities in the sense of getting some kind of lower and upper estimates. Now obviously there are some problems that current LLMs still can't solve: we know for theoretical reasons that there always will be, but we also know this more constructively because if someone had gotten their favorite model to solve one of the millennium problems, we would almost certainly have heard them boasting all over the Internet as if they had been part of the achievement: so by a “dog that didn't bark” https://en.wikipedia.org/wiki/The Adventure of Silver Blaze argument, I conclude that such problems are currently still out of reach of even the best LLMs that OpenAI, Anthropic &co can make internally. Still, I suspect that the upper bound can be improved beyond “ChatGPT can't solve the Riemann hypothesis”. In fact, I suspect that problems still exist that human mathematicians can solve with relative ease and LLMs can't, even if they maybe have to be constructed in a somewhat “adversarial” way: not so long ago, after all, some of the best models failed to accurately count the number of ‘r’s in “strawberry” or answer the question of whether there is a seahorse emoji https://www.youtube.com/watch?v=W2xZxYaGlfs : these examples can be attributed to oddities in the way LLMs work, but I see no reason why there shouldn't be such oddities within mathematics e.g., maybe a conclusipn that can be easily seen by drawing a picture can be very hard for an LLM to reach . Anyway: Question: Are there known examples of math problems that even the best current AI models fail to solve in a reasonable amount of time or solve incorrectly without human assistance, and which average specialist-in-the-field human mathematicians can solve?