cd /news/artificial-intelligence/arise-rl-agentic-rubric-grounded-ite… · home topics artificial-intelligence article
[ARTICLE · art-118674] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

Researchers propose ARISE-RL, a self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution to train open-ended agents via reinforcement learning, addressing the lack of verifiable gold answers and scalable rubrics. The framework introduces Reward-Gated Self-Evolution Distillation (RG-SED) to selectively distill memory-augmented policy variants, and presents ECR-Bench, an expert-calibrated rubric benchmark suite for single-tool deep research and multi-tool travel planning. Experiments show ARISE-RL achieves state-of-the-art performance across all evaluated benchmarks.

read1 min views1 publishedSep 2, 2026

arXiv:2609.01058v1 Announce Type: new Abstract: Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution. The Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks aligned with the Solver's evolving capability boundary. The Solver, in turn, learns from fine-grained rubric satisfaction signals through multi-step reasoning and tool use. We further introduce Reward-Gated Self-Evolution Distillation (RG-SED), which selectively distills a memory-augmented variant of the same policy back into itself only when the memory yields empirical reward improvement, thereby reducing distribution mismatch and avoiding blind imitation of noisy guidance. Finally, to support rigorous evaluation, we present ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning. Extensive experiments demonstrate that ARISE-RL consistently achieves robust and stable overall state-of-the-art performance across all evaluated benchmarks.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arise-rl 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/arise-rl-agentic-rub…] indexed:0 read:1min 2026-09-02 ·