# The Leiden Declaration on AI and Math just dropped — has anyone

> Source: <https://promptcube3.com/en/news/7246/>
> Published: 2026-08-22 01:08:35+00:00

# The Leiden Declaration on AI and Math just dropped — has anyone

What grabbed me: they're pushing for a "mathematical Turing test" variant where the evaluator is another mathematician, not an LLM judge. The protocol involves three phases — problem formulation, collaborative exploration, and final proof — with human experts rating whether the AI contributed meaningful mathematical insight at each stage. No multiple choice, no formal verification checkmarks alone.

Has anyone seen the evaluation rubric they proposed? Appendix B lists criteria like "identifies necessary lemmas without prompting" and "recognizes dead ends and backtracks strategically." That second one feels huge — most models just hallucinate forward. If an agent can genuinely say "this approach won't work because X" and pivot, that's closer to how mathematicians actually think.

Also curious about their stance on formalization. They acknowledge Lean/Isabelle as necessary infrastructure but explicitly reject "formalization coverage" as a proxy for understanding. Their example: an AI that formalizes a known proof in Lean gets zero credit for mathematical novelty. Fair, but then how do you quantify the "insight" metric without human graders? The declaration calls for a standing committee of Fields medalists and IMO coaches — seems... optimistic for funding.

One concrete thing I'll test this weekend: their suggested "minimal viable benchmark" — 50 problems from recent Putnam exams, none in training data, evaluated by two independent graders using their rubric. Small enough to run manually, representative enough to matter. If anyone wants to collaborate on grading, DM me.

The declaration site has a sign-up for the first evaluation round in Q2. Might be worth joining just to see how the rubric holds up in practice.

[Next Proliferate unifies Claude Code, Codex →](/en/news/7241/)
