cd /news/artificial-intelligence/counterexample-generation-via-per-th… · home › topics › artificial-intelligence › article
[ARTICLE · art-145165] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs

Training Qwen3-4B with counterexample-only supervised fine-tuning collapsed true-theorem recognition from 0.27 to 0.00, while reinforcement learning with a sparse outcome-only reward repaired the collapse and exceeded the base model at 0.66, according to an arXiv paper releasing SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures each paired with an executable per-theorem Python verifier. The collapse replicated across four seeds and on Gemma-3-4B, and the resulting 4B model outperformed every evaluated 7B open-weights math specialist, stayed competitive with six frontier commercial APIs, and transferred under unchanged prompting to GSM8K, MATH-500 and MMLU-college-math. A human audit of 177 verifier decisions found 97.7% accuracy, and code, data, verifier modules and annotations are available at https://github.com/ce-rlvr/SymCE.

by read1 min views1 publishedOct 5, 2026

arXiv:2610.02444v1 Announce Type: new Abstract: Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python verifier, and release SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures, each paired with executable verifiers. The verifier also serves as the reward function, making SymCE a training environment. Training Qwen3-4B with SFT followed by GRPO under this oracle reveals an imitation trap: counterexample-only SFT collapses true-theorem recognition from 0.27 to 0.00, while RLVR with a sparse outcome-only reward repairs this and exceeds the base, to 0.66. The collapse replicates across four seeds and on Gemma-3-4B. Sparse and dense rewards yield statistically indistinguishable in-domain success yet diverge by 33 points on a held-out calibration probe, a dissociation we trace to the partial-credit term. Our 4B model outperforms every evaluated 7B open-weights math specialist, remains competitive with six frontier commercial APIs, and transfers under unchanged prompting to GSM8K, MATH-500 and MMLU-college-math. A human audit of 177 verifier decisions finds 97.7% accuracy. Code, data, verifier modules and annotations: https://github.com/ce-rlvr/SymCE.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @qwen3-4b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/counterexample-gener…] indexed:0 read:1min 2026-10-05 · —