04:00
2026-10-05
arxiv.org
artificial-intelligence
Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs
Training Qwen3-4B with counterexample-only supervised fine-tuning collapsed true-theorem recognition from 0.27 to 0.00, while reinforcement learning with a sparse outcome-only reward repaired the coll…