Assessment of open AI math results In a social media post, an unnamed user reported that OpenAI's GPT-5.6 Sol Pro and Fable 5 Max, two AI models, classified results from Epoch AI's OpenMath benchmark using its rubric, with both models agreeing that result #3 is a 'Breakthrough' and at least 7 results are 'Major Advancements'. The models differed on result #7, with Fable 5 Max labeling it 'Solid Result' and Sol Pro labeling it 'Major Advance', while Fable 5 Max considered results #1, #4, and #9 as 'Borderline Breakthrough'. The post also noted that Epoch AI's rubric errs on the conservative side when multiple tiers seem plausible. It's hard for an ordinary person to understand the complexity of these tasks. I'm no mathematician, and I don't see a difference between e.g., results 3 and 10. So I had GPT-5.6 Sol Pro and Fable 5 Max classify these using @EpochAIResearch https://x.com/EpochAIResearch OpenMath's rubric: — "Solid Result": A strong researcher in the area would be happy if their median output addressed problems of this caliber. Still, the problem would probably not get much engagement outside of its subfield. — "Major Advance": The median person working in a broad area of mathematics on the scale of number theory or graph theory would take note, and would likely make the time to understand at least the outline of the solution. — "Breakthrough": The median mathematician would want to know about this result, even if it was outside their area. It would be a candidate for one of the best results of the year in all of mathematics === Both Fable and Sol agree 3 is a Breakthrough which explains why @SebastienBubeck https://x.com/SebastienBubeck opens his tweet with it . They also agree that at least 7 are Major Advancements. Fable thinks 7 is just a Solid Result, while Sol assigns the "Major Advance" label. What's also interesting is that the official Epoch.AI http://Epoch.AI rubrics say this: When multiple tiers seemed plausible for a problem, we erred in the conservative direction. It would be disappointing to downgrade a problem’s notability after it was solved, whereas we can always highlight any unexpectedly interesting elements of a solution. And Fable 5 thinks that at least 3 of the results are "Borderline Breakthrough" 1, 4, and 9 . @AcerFur https://x.com/AcerFur any thoughts on thisAn internal version of Astra,