{"slug": "counterexample-generation-via-per-theorem-symbolic-verifiers-when-imitation-and", "title": "Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs", "summary": "Training Qwen3-4B with counterexample-only supervised fine-tuning collapsed true-theorem recognition from 0.27 to 0.00, while reinforcement learning with a sparse outcome-only reward repaired the collapse and exceeded the base model at 0.66, according to an arXiv paper releasing SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures each paired with an executable per-theorem Python verifier. The collapse replicated across four seeds and on Gemma-3-4B, and the resulting 4B model outperformed every evaluated 7B open-weights math specialist, stayed competitive with six frontier commercial APIs, and transferred under unchanged prompting to GSM8K, MATH-500 and MMLU-college-math. A human audit of 177 verifier decisions found 97.7% accuracy, and code, data, verifier modules and annotations are available at https://github.com/ce-rlvr/SymCE.", "body_md": "arXiv:2610.02444v1 Announce Type: new \nAbstract: Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python verifier, and release SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures, each paired with executable verifiers. The verifier also serves as the reward function, making SymCE a training environment. Training Qwen3-4B with SFT followed by GRPO under this oracle reveals an imitation trap: counterexample-only SFT collapses true-theorem recognition from 0.27 to 0.00, while RLVR with a sparse outcome-only reward repairs this and exceeds the base, to 0.66. The collapse replicates across four seeds and on Gemma-3-4B. Sparse and dense rewards yield statistically indistinguishable in-domain success yet diverge by 33 points on a held-out calibration probe, a dissociation we trace to the partial-credit term. Our 4B model outperforms every evaluated 7B open-weights math specialist, remains competitive with six frontier commercial APIs, and transfers under unchanged prompting to GSM8K, MATH-500 and MMLU-college-math. A human audit of 177 verifier decisions finds 97.7% accuracy. Code, data, verifier modules and annotations: https://github.com/ce-rlvr/SymCE.", "url": "https://wpnews.pro/news/counterexample-generation-via-per-theorem-symbolic-verifiers-when-imitation-and", "canonical_source": "https://arxiv.org/abs/2610.02444", "published_at": "2026-10-05 04:00:00+00:00", "updated_at": "2026-10-05 04:14:44.184578+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["Qwen3-4B", "Gemma-3-4B", "SymCE", "GSM8K", "MATH-500", "MMLU-college-math", "arXiv", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/counterexample-generation-via-per-theorem-symbolic-verifiers-when-imitation-and", "markdown": "https://wpnews.pro/news/counterexample-generation-via-per-theorem-symbolic-verifiers-when-imitation-and.md", "text": "https://wpnews.pro/news/counterexample-generation-via-per-theorem-symbolic-verifiers-when-imitation-and.txt", "jsonld": "https://wpnews.pro/news/counterexample-generation-via-per-theorem-symbolic-verifiers-when-imitation-and.jsonld"}}