Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation

wpnews.pro

cd /news/artificial-intelligence/pythagoras-prover-advancing-efficien… · home › topics › artificial-intelligence › article

[ARTICLE · art-24795] src=arxiv.org ↗ pub=2026-06-12T04:00Z topic=artificial-intelligence verified=true sentiment=↑ positive

Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation

Researchers introduced Pythagoras-Prover, a family of compute-efficient Lean theorem provers that achieve state-of-the-art performance with significantly fewer parameters. The 4B-parameter model surpassed DeepSeek-Prover-V2-671B on the MiniF2F-Test benchmark with 86.1% accuracy using 167x fewer parameters, while the 32B model set an open-source record at 93.0% and solved 93 PutnamBench problems. The system's efficiency stems from a curriculum training approach using stratified verified proofs and Augmented Lean Formalisation, which expands scarce training data through mutated formal statements.

read1 min publishedJun 12, 2026

arXiv:2606.12594v1 Announce Type: new Abstract: Modern Lean theorem provers achieve strong performance only with substantial training and inference compute, driven in part by scarce verified proof data and the long reasoning traces of formal proof search, making both supervised fine-tuning (SFT) and sampling expensive. We introduce Pythagoras-Prover, a compute-efficient open-source family of Lean theorem provers built for practical compute budgets. The family spans two generation paradigms: autoregressive models at 4B and 32B parameters, and a first proof-of-concept diffusion-based prover (4B) that iteratively refines Lean proofs at inference time. For training efficiency, we build a Lean-verified corpus stratified into easy, medium, and hard problems for curriculum SFT, so models acquire proof skills progressively from shorter, simpler proofs to longer, harder ones. During SFT, a dynamic proof-reasoning filtering scheme preserves informative proof traces while keeping each instance within an 8k-token context budget. We also introduce Augmented Lean Formalisation (ALF), which expands scarce verified corpora into variants of formal statements, populated via self-distillation for extra training signal without formally verifying every mutated instance. By perturbing known problems while preserving their formal character, ALF reduces reliance on any statement's surface form. Empirically, Pythagoras-Prover-4B surpasses DeepSeek-Prover-V2-671B at pass@32 on MiniF2F-Test (86.1% vs 82.4%) with ~167x fewer parameters, while Pythagoras-Prover-32B sets the open-source state of the art at 93.0% on MiniF2F-Test and solves 93 of 672 PutnamBench problems. We release MiniF2F-ALF, an ALF-mutated contamination-sensitive benchmark on which every evaluated model loses accuracy; here our 32B remains strongest and our 4B matches the prior state of the art, Goedel-Prover-V2-32B.

source & further reading

arxiv.org — original article

~/api · this article 200

$curl api.wpnews.pro/v1/news/pythagoras-prover-advanc…

Read original on arxiv.org → arxiv.org/abs/2606.12594

mentioned entities

Pythagoras-Prover

DeepSeek-Prover-V2

MiniF2F-Test

Lean

metadata

slugpythagoras-prover-advancing-efficient-formal-proving-via-augmented-lean

topic#artificial-intelligence

secondary3 topics

sentimentpositive

langen

canonicalarxiv.org

navigation

← prevLinear Coding Sessions

next →Can KKR Outmaneuver One of the B…

── more in #artificial-intelligence 4 stories · sorted by recency

lesswrong.com · 13 Jun · #artificial-intelligence

SFT Drives Gemini’s Safety Properties

lesswrong.com · 13 Jun · #artificial-intelligence

The term “AGI” is almost useless at this point [Linkpost]

dev.to · 13 Jun · #artificial-intelligence

Released larkos 0.3

signal-memo.com · 13 Jun · #artificial-intelligence

AI Benchmarks Are Starting to Look Like Emissions Tests

sponsored brought to you by zahid.host 4,200+ EU-deployed projects

reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main

→ Live at https://your-agent.zahid.host ✓

Get free account → Pricing

from €0/mo · no card required