{"slug": "multilingual-verifier-bias-in-rlvr-benchmark-rollout-diagnosis-and-the-cross", "title": "Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck", "summary": "A new arXiv study (2608.20362v1) finds that exact-match verifiers in reinforcement learning with verifiable rewards (RLVR) introduce language-dependent false-negative reward noise in multilingual settings, with Qwen3-8B showing a false-negative rate of 0.642 on Japanese answers versus 0.122 on English and 0.073 on Chinese on MGSM rollouts. The authors propose a reusable audit protocol and show that a target-local aggregation rule closes 55-78% of the average selection gap on MGSM250, but over 95% of repairs require cross-lingual support, highlighting a cross-lingual selection bottleneck.", "body_md": "arXiv:2608.20362v1 Announce Type: new\nAbstract: Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a language-neutral reward function. We show that this assumption fails in multilingual settings: an exact-match verifier turns format and script variation into language-dependent false-negative reward noise. We introduce a reusable protocol for auditing multilingual RLVR rewards: a verifier-robustness suite, a rollout-diagnosis procedure, and language-conditioned reward-error metrics for Japanese, English, and Chinese answers. On MGSM rollouts with k=8, the exact-match proxy rejects trusted-correct answers at sharply different rates by language across Qwen3-4B, Qwen3-8B, and Llama-3.1-8B-Instruct; for Qwen3-8B, the false-negative rate reaches 0.642 on JP against 0.122 on EN and 0.073 on CN. A plain-numeric probe localizes the mechanism to the final-answer interface: an interface model drives reward-error VLB to zero while the residual accuracy gap is unchanged. We then expose a cross-lingual selection bottleneck: on MGSM250 rollouts, a target-local aggregation rule using no trusted labels closes 55-78% of the average selection gap, and over 95% of repairs require genuine cross-lingual support. The bottleneck replicates on a 483-problem MATH-500 set. A controlled training audit shows that rule-GRPO raises trusted accuracy while the reward-error VLB stays high. The unifying message is operational: multilingual RLVR rewards should be audited by language and by answer interface before they are optimized.", "url": "https://wpnews.pro/news/multilingual-verifier-bias-in-rlvr-benchmark-rollout-diagnosis-and-the-cross", "canonical_source": "https://arxiv.org/abs/2608.20362", "published_at": "2026-08-24 04:00:00+00:00", "updated_at": "2026-08-24 04:14:36.272366+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["arXiv", "Qwen3-4B", "Qwen3-8B", "Llama-3.1-8B-Instruct", "MGSM", "MATH-500"], "alternates": {"html": "https://wpnews.pro/news/multilingual-verifier-bias-in-rlvr-benchmark-rollout-diagnosis-and-the-cross", "markdown": "https://wpnews.pro/news/multilingual-verifier-bias-in-rlvr-benchmark-rollout-diagnosis-and-the-cross.md", "text": "https://wpnews.pro/news/multilingual-verifier-bias-in-rlvr-benchmark-rollout-diagnosis-and-the-cross.txt", "jsonld": "https://wpnews.pro/news/multilingual-verifier-bias-in-rlvr-benchmark-rollout-diagnosis-and-the-cross.jsonld"}}