{"slug": "repro-proof-verified-benchmark-rewriting-for-reliable-evaluation-of-llm-problem", "title": "RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving", "summary": "Researchers introduced Proof-Verified Benchmark Rewriting (RePro), the first framework integrating Lean-oriented neural automated theorem provers into benchmark rewriting, ensuring rewritten math problems and answers are verified by Lean proofs. On GSM8K and MATH, RePro's retained instances achieved 100% well-definedness, feasibility, and answer correctness, while existing methods produced invalid or incorrect instances. Several models showed accuracy drops on proof-verified benchmarks, indicating sensitivity to surface variations and possible memorization effects.", "body_md": "arXiv:2609.00062v1 Announce Type: new\nAbstract: Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, existing methods lack guarantees of problem validity and answer correctness. We propose Proof-Verified Benchmark Rewriting (RePro), the first framework to integrate Lean-oriented neural automated theorem provers (ATPs) into benchmark rewriting, which rewrites problems and regenerates answers with correctness ensured by Lean-verified proofs. Experiments on GSM8K and MATH show that RePro's retained rewritten instances achieve 100% well-definedness, feasibility, and answer correctness, while existing methods still produce invalid or incorrect instances. Moreover, several models exhibit accuracy drops on proof-verified rewritten benchmarks, suggesting that their performance is sensitive to surface-level and structural variations and may partly reflect memorization effects. Our source code and data are available at https://github.com/AI4Engi/RePro.", "url": "https://wpnews.pro/news/repro-proof-verified-benchmark-rewriting-for-reliable-evaluation-of-llm-problem", "canonical_source": "https://arxiv.org/abs/2609.00062", "published_at": "2026-09-02 04:00:00+00:00", "updated_at": "2026-09-02 04:25:52.852797+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["RePro", "Lean", "GSM8K", "MATH", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/repro-proof-verified-benchmark-rewriting-for-reliable-evaluation-of-llm-problem", "markdown": "https://wpnews.pro/news/repro-proof-verified-benchmark-rewriting-for-reliable-evaluation-of-llm-problem.md", "text": "https://wpnews.pro/news/repro-proof-verified-benchmark-rewriting-for-reliable-evaluation-of-llm-problem.txt", "jsonld": "https://wpnews.pro/news/repro-proof-verified-benchmark-rewriting-for-reliable-evaluation-of-llm-problem.jsonld"}}