{"slug": "the-leiden-declaration-on-ai-and-math-just-dropped-has-anyone", "title": "The Leiden Declaration on AI and Math just dropped — has anyone", "summary": "The Leiden Declaration on AI and Mathematics proposes a 'mathematical Turing test' where AI systems are evaluated by human mathematicians across three phases—problem formulation, collaborative exploration, and final proof—rejecting formalization coverage as a proxy for understanding. The declaration includes a rubric with criteria such as 'identifies necessary lemmas without prompting' and 'recognizes dead ends and backtracks strategically,' and calls for a standing committee of Fields medalists and IMO coaches. A minimal viable benchmark of 50 Putnam exam problems is suggested for testing, with the first evaluation round planned for Q2.", "body_md": "# The Leiden Declaration on AI and Math just dropped — has anyone\n\nWhat grabbed me: they're pushing for a \"mathematical Turing test\" variant where the evaluator is another mathematician, not an LLM judge. The protocol involves three phases — problem formulation, collaborative exploration, and final proof — with human experts rating whether the AI contributed meaningful mathematical insight at each stage. No multiple choice, no formal verification checkmarks alone.\n\nHas anyone seen the evaluation rubric they proposed? Appendix B lists criteria like \"identifies necessary lemmas without prompting\" and \"recognizes dead ends and backtracks strategically.\" That second one feels huge — most models just hallucinate forward. If an agent can genuinely say \"this approach won't work because X\" and pivot, that's closer to how mathematicians actually think.\n\nAlso curious about their stance on formalization. They acknowledge Lean/Isabelle as necessary infrastructure but explicitly reject \"formalization coverage\" as a proxy for understanding. Their example: an AI that formalizes a known proof in Lean gets zero credit for mathematical novelty. Fair, but then how do you quantify the \"insight\" metric without human graders? The declaration calls for a standing committee of Fields medalists and IMO coaches — seems... optimistic for funding.\n\nOne concrete thing I'll test this weekend: their suggested \"minimal viable benchmark\" — 50 problems from recent Putnam exams, none in training data, evaluated by two independent graders using their rubric. Small enough to run manually, representative enough to matter. If anyone wants to collaborate on grading, DM me.\n\nThe declaration site has a sign-up for the first evaluation round in Q2. Might be worth joining just to see how the rubric holds up in practice.\n\n[Next Proliferate unifies Claude Code, Codex →](/en/news/7241/)", "url": "https://wpnews.pro/news/the-leiden-declaration-on-ai-and-math-just-dropped-has-anyone", "canonical_source": "https://promptcube3.com/en/news/7246/", "published_at": "2026-08-22 01:08:35+00:00", "updated_at": "2026-08-22 01:12:49.482642+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-ethics"], "entities": ["Leiden Declaration", "Lean", "Isabelle", "Putnam", "Fields Medal", "IMO"], "alternates": {"html": "https://wpnews.pro/news/the-leiden-declaration-on-ai-and-math-just-dropped-has-anyone", "markdown": "https://wpnews.pro/news/the-leiden-declaration-on-ai-and-math-just-dropped-has-anyone.md", "text": "https://wpnews.pro/news/the-leiden-declaration-on-ai-and-math-just-dropped-has-anyone.txt", "jsonld": "https://wpnews.pro/news/the-leiden-declaration-on-ai-and-math-just-dropped-has-anyone.jsonld"}}