cd /news/artificial-intelligence/multilingual-verifier-bias-in-rlvr-b… · home topics artificial-intelligence article
[ARTICLE · art-108254] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck

A new arXiv study (2608.20362v1) finds that exact-match verifiers in reinforcement learning with verifiable rewards (RLVR) introduce language-dependent false-negative reward noise in multilingual settings, with Qwen3-8B showing a false-negative rate of 0.642 on Japanese answers versus 0.122 on English and 0.073 on Chinese on MGSM rollouts. The authors propose a reusable audit protocol and show that a target-local aggregation rule closes 55-78% of the average selection gap on MGSM250, but over 95% of repairs require cross-lingual support, highlighting a cross-lingual selection bottleneck.

read1 min views2 publishedAug 24, 2026

arXiv:2608.20362v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a language-neutral reward function. We show that this assumption fails in multilingual settings: an exact-match verifier turns format and script variation into language-dependent false-negative reward noise. We introduce a reusable protocol for auditing multilingual RLVR rewards: a verifier-robustness suite, a rollout-diagnosis procedure, and language-conditioned reward-error metrics for Japanese, English, and Chinese answers. On MGSM rollouts with k=8, the exact-match proxy rejects trusted-correct answers at sharply different rates by language across Qwen3-4B, Qwen3-8B, and Llama-3.1-8B-Instruct; for Qwen3-8B, the false-negative rate reaches 0.642 on JP against 0.122 on EN and 0.073 on CN. A plain-numeric probe localizes the mechanism to the final-answer interface: an interface model drives reward-error VLB to zero while the residual accuracy gap is unchanged. We then expose a cross-lingual selection bottleneck: on MGSM250 rollouts, a target-local aggregation rule using no trusted labels closes 55-78% of the average selection gap, and over 95% of repairs require genuine cross-lingual support. The bottleneck replicates on a 483-problem MATH-500 set. A controlled training audit shows that rule-GRPO raises trusted accuracy while the reward-error VLB stays high. The unifying message is operational: multilingual RLVR rewards should be audited by language and by answer interface before they are optimized.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/multilingual-verifie…] indexed:0 read:1min 2026-08-24 ·