{"slug": "leanpolish-verified-supervision-for-lean-proof-compression", "title": "LeanPolish: Verified Supervision for Lean Proof Compression", "summary": "A symbolic Lean 4 pipeline called LeanPolish released 33,402 accepted local edits and 65,596 same-state failed attempts to study what language models learn from verified proof-edit supervision, according to the arXiv paper. Continuing menu evaluation beyond the first success removed an ordering shortcut, letting a trained ranker select the best candidate on 70.1% of evaluated held-out states versus 36.9% for the strongest frozen baseline, and iterating the symbolic pass raised miniF2F savings from 19.7% to 27.5%. Fine-tuning raised verified token reduction from 2.8% to 5.5% on 19 PutnamBench proofs, while matched frozen-model controls showed verified neural editing gains need not come from training.", "body_md": "arXiv:2609.38384v1 Announce Type: new \nAbstract: Verified proof edits offer a natural source of supervision for improving language-model-generated Lean proofs. Yet verification establishes that an edit is correct, not that its training signal is free of search artifacts. We introduce LeanPolish, a symbolic Lean 4 pipeline that releases 33,402 accepted local edits and 65,596 same-state failed attempts, and use it to study what models learn from this supervision. First-success search admits a goal-independent rule with perfect ranking accuracy; teacher-selected evaluation sites also reward trivial deletions. Continuing menu evaluation beyond the first success removes the ordering shortcut: a trained ranker selects the best candidate on 70.1% of evaluated held-out states, versus 36.9% for the strongest frozen baseline. For compression, iterating the symbolic pass raises miniF2F savings from 19.7% to 27.5%, exceeding the neural hybrids we test there. Verified neural editing helps on other proof sources, but matched frozen-model controls show that its gains need not come from training. The supervision does improve whole-proof rewriting: fine-tuning raises verified token reduction from 2.8% to 5.5% on 19 PutnamBench proofs. Together, the released edits, complete candidate pools, and controlled evaluations separate learning to imitate a search policy from improving on that search. They provide a reproducible basis for studying proof improvement while keeping correctness, compression, and edit policy distinct.", "url": "https://wpnews.pro/news/leanpolish-verified-supervision-for-lean-proof-compression", "canonical_source": "https://arxiv.org/abs/2609.38384", "published_at": "2026-10-02 04:00:00+00:00", "updated_at": "2026-10-02 04:16:38.160957+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["LeanPolish", "Lean 4", "miniF2F", "PutnamBench", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/leanpolish-verified-supervision-for-lean-proof-compression", "markdown": "https://wpnews.pro/news/leanpolish-verified-supervision-for-lean-proof-compression.md", "text": "https://wpnews.pro/news/leanpolish-verified-supervision-for-lean-proof-compression.txt", "jsonld": "https://wpnews.pro/news/leanpolish-verified-supervision-for-lean-proof-compression.jsonld"}}