Examples of problems that current AI models fail to solve but people can A MathOverflow user is soliciting examples of mathematical problems that current AI models fail to solve but that an averagely competent mathematician can solve, aiming to counter the observation bias in reporting AI successes. The user clarifies that failures need not be reproducible and that the goal is to gather anecdotal evidence to better understand LLM limitations. Important preliminary note added 2026-08-19T23:58+02:00 : The comments below show and the downvotes probably suggest that various people have interpreted my question as some combination of ⓐ representing problem-solving as a “humans versus AI” game, ⓑ wishing to either demonstrate human superiority in this matter or on the contrary to help improve AIs, and/or lastly ⓒ reducing mathematical practice to problem-solving. These three interpretations are all incorrect I can point to lengthy threads of mine on other social media to prove this, but I don't think MathOverflow is the appropriate place to discuss this . The point of the question is to get a slightly less observation-biased understanding of what LLMs, in their current status, can and cannot do: we have enormous amounts of data about the “can”, so I'm trying to get some grasp on the “cannot” and comparison with human beings is a natural point of reference because that is how we conceive of “difficulty” : this is in no way to be taken to be taken as any subjective judgement on my part about humans-versus-AI, or problem-solving, or anything of the sort. Another clarification added 2026-08-20T13:42+02:00 : Several people have worried in the comment that “AI fails to solve problem $X$” is not a reproducible fact. So, to be clear, I am not asking for examples where the LLM fails reproducibly and systematically to solve the problem — if it fails under reasonable conditions i.e., at least one reasonably state-of-the-art model fails when given a reasonable amount of time , this qualifies as an example: please exercise your own judgment in what “reasonable”. The same applies to the counterpart: I am not requiring that each and every person with knowledge in the field is able to solve the problem, only that we have reasonable suspicion that an averagely competent mathematician can solve it — again, exercise your own good faith judgment in deciding what this means. Something like: “I tried asking model to prove lemma , it failed, and I later realized it wasn't hard at all” is completely acceptable. Background and motivation: There is now a long list of problems https://mathoverflow.net/questions/502120/examples-for-the-use-of-ai-and-especially-llms-in-major-mathematical-development , including many long-standing conjectures, that have been solved autonomously with current AI models. The problem with this sort of list is that it conveys the idea that LLMs are now far better than human beings at problem-solving because they are able to solve many problems that people couldn't: while the conclusion might be correct, the reasoning certainly is not: it might also be the case that problem solving difficulty for LLMs and human brains is not well-correlated, and selecting problems hard for human beings still leaves a lot of problems that are easy for AI selection bias . Underlying this is also the huge observation bias that “AI model solves long-standing conjecture” is considered newsworthy and attentionworthy whereas “AI model fails to solve problem that turns out to be not that difficult after all” won't get reported. There is the additional difficulty that any problem previously solved by people is probably already in the AI models' training set. I am aware of the First Proof project https://1stproof.org/ which seeks to confront this observation bias in a more scientific way, but I think it is still interesting to gather anecdotal evidence to put the positive AI news into perspective. Anyway, if we are to navigate this brave new world that has such AI in't, we need to understand its abilities in the sense of getting some kind of lower and upper estimates. Now obviously there are some problems that current LLMs still can't solve: we know for theoretical reasons that there always will be, but we also know this more constructively because if someone had gotten their favorite model to solve one of the millennium problems, we would almost certainly have heard them boasting all over the Internet as if they had been part of the achievement: so by a “dog that didn't bark” https://en.wikipedia.org/wiki/The Adventure of Silver Blaze argument, I conclude that such problems are currently still out of reach of even the best LLMs that OpenAI, Anthropic &co can make internally. Still, I suspect that the upper bound can be improved beyond “ChatGPT can't solve the Riemann hypothesis”. In fact, I suspect that problems still exist that mathematicians can solve with relative ease and LLMs can't, even if they maybe have to be constructed in a somewhat “adversarial” way: not so long ago, after all, some of the best models failed to accurately count the number of ‘r’s in “strawberry” or answer the question of whether there is a seahorse emoji https://www.youtube.com/watch?v=W2xZxYaGlfs : these examples can be attributed to oddities in the way LLMs work, but I see no reason why there shouldn't be such oddities within mathematics e.g., maybe a conclusipn that can be easily seen by drawing a picture can be very hard for an LLM to reach . Anyway: Question: Are there known examples of math problems that even the best current AI models fail to solve in a reasonable amount of time or solve incorrectly without human assistance, and which average specialist-in-the-field mathematicians can solve without AI assistance?