cd /news/artificial-intelligence/the-leiden-declaration-on-ai-and-mat… · home topics artificial-intelligence article
[ARTICLE · art-106695] src=promptcube3.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

The Leiden Declaration on AI and Math just dropped — has anyone

The Leiden Declaration on AI and Mathematics proposes a 'mathematical Turing test' where AI systems are evaluated by human mathematicians across three phases—problem formulation, collaborative exploration, and final proof—rejecting formalization coverage as a proxy for understanding. The declaration includes a rubric with criteria such as 'identifies necessary lemmas without prompting' and 'recognizes dead ends and backtracks strategically,' and calls for a standing committee of Fields medalists and IMO coaches. A minimal viable benchmark of 50 Putnam exam problems is suggested for testing, with the first evaluation round planned for Q2.

read1 min views1 publishedAug 22, 2026
The Leiden Declaration on AI and Math just dropped — has anyone
Image: Promptcube3 (auto-discovered)

What grabbed me: they're pushing for a "mathematical Turing test" variant where the evaluator is another mathematician, not an LLM judge. The protocol involves three phases — problem formulation, collaborative exploration, and final proof — with human experts rating whether the AI contributed meaningful mathematical insight at each stage. No multiple choice, no formal verification checkmarks alone.

Has anyone seen the evaluation rubric they proposed? Appendix B lists criteria like "identifies necessary lemmas without prompting" and "recognizes dead ends and backtracks strategically." That second one feels huge — most models just hallucinate forward. If an agent can genuinely say "this approach won't work because X" and pivot, that's closer to how mathematicians actually think.

Also curious about their stance on formalization. They acknowledge Lean/Isabelle as necessary infrastructure but explicitly reject "formalization coverage" as a proxy for understanding. Their example: an AI that formalizes a known proof in Lean gets zero credit for mathematical novelty. Fair, but then how do you quantify the "insight" metric without human graders? The declaration calls for a standing committee of Fields medalists and IMO coaches — seems... optimistic for funding.

One concrete thing I'll test this weekend: their suggested "minimal viable benchmark" — 50 problems from recent Putnam exams, none in training data, evaluated by two independent graders using their rubric. Small enough to run manually, representative enough to matter. If anyone wants to collaborate on grading, DM me.

The declaration site has a sign-up for the first evaluation round in Q2. Might be worth joining just to see how the rubric holds up in practice.

Next Proliferate unifies Claude Code, Codex →

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @leiden declaration 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-leiden-declarati…] indexed:0 read:1min 2026-08-22 ·