arXiv:2610.11118v1 Announce Type: new Abstract: The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge. We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature. Each problem supplies the research context, assumptions, and prior progress needed to investigate the question. We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims. Four evaluator models independently assess the correctness, completeness, and degree of progress of each submission without reference solutions. Across seven evaluated configurations, GPT-6-Astra achieves the highest mean judged solve rate of 14.0%, compared with 5.5-6.7% for the evaluated full-size open models and 2.4-3.7% for Flash models. Case comparisons connect stronger outcomes to changes in problem representation, general arguments that extend beyond finite evidence, and proofs of the steps needed to complete a solution. By grounding evaluation in questions arising from the research literature, OpenProblemBench provides a setting for investigating the capabilities and limitations of AI as a contributor to foundational theoretical science.
OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences
OpenProblemBench, a new benchmark of 82 unresolved problems from the mathematics and theoretical physics literature, reports that GPT-6-Astra achieved the highest mean judged solve rate at 14.0% across seven evaluated configurations, according to the arXiv paper 2610.11118v1. Full-size open models scored 5.5-6.7% and Flash models 2.4-3.7%, with four evaluator models independently judging correctness, completeness, and degree of progress without reference solutions. The benchmark supplies research context, assumptions, and prior progress for each problem to test AI as a contributor to foundational theoretical science.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.