Beyond Solver Verdicts: Generative Reward Models for Autoformalization Researchers formalized a vulnerability they call Verdict-Preserving-Unfaithfulness (VPU), in which mathematical solvers in neurosymbolic systems verify reasoning correctness but cannot detect whether a formal translation maintains strict reference-equivalence to a designated formalization. The work proposes generative reward models for autoformalization to address the gap left by solver verdicts alone. Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally blind to whether a formal translation maintains strict reference-equivalence to a designated formalization. We formalize this vulnerability as Verdict-Preserving-Unfaithfulness VPU :