cd /news/artificial-intelligence/split-rl-staged-perception-language-… · home › topics › artificial-intelligence › article
[ARTICLE · art-148010] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

SPLIT-RL, a staged post-training approach that trains visual reasoning and language reasoning in disjoint phases, improves average accuracy over GRPO by 1.4 to 6.1 points across Qwen3-VL models from 2B to 30B-A3B and InternVL3.5-8B, according to the arXiv paper 2610.10889v1. The method also introduces Claim-Level Advantage (CLA-GRPO), which decomposes visual-reasoning-phase rollouts into atomic visual claims and assigns fine-grained advantage at the claim level based on visual-type group formation. An oracle-based diagnostic showed answer-only GRPO leaves perception unchanged, whereas SPLIT-RL improves both visual and language reasoning, with the trained policy evaluated using a single chain-of-thought call at inference time.

by read1 min views3 publishedOct 9, 2026

arXiv:2610.10889v1 Announce Type: new Abstract: Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR). Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer reward. This gives every CoT token the same sequence-level advantage, failing to distinguish capability specific errors. We propose SPLIT-RL, a staged post-training approach that trains VR and LR in disjoint phases. Because a group's rollouts differ along one capability at a time, the group-relative advantage isolates it, and each phase is optimized using phase-specific reward. We further introduce Claim-Level Advantage (CLA-GRPO), which decomposes VR-phase rollouts into atomic visual claims and provides a fine-grained advantage at claim level based on visual-type group formation. Although trained in two phases, trained policy is evaluated like GRPO model, with a single CoT call at inference time. Under this protocol, SPLIT-RL improves average accuracy over GRPO by 1.4-6.1 points across Qwen3-VL models from 2B to 30B-A3B and InternVL3.5-8B. Evaluating each capability using an oracle based diagnostic shows that answer-only GRPO leaves perception unchanged, whereas SPLIT-RL improves both VR and LR.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @split-rl 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/split-rl-staged-perc…] indexed:0 read:1min 2026-10-09 · —