cd /news/artificial-intelligence/rised-rubrics-for-agentic-multi-envi… · home › topics › artificial-intelligence › article
[ARTICLE · art-146487] src=machinelearning.apple.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

Researchers Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low and Manjot Bilkhu published RISED, a method that repurposes rubrics — textual descriptions of rollout behaviours tagged by an LLM judge — to guide both online data selection and policy supervision when training a single LLM agent across diverse interactive environments. RISED uses positive rubrics as privileged context for an on-policy self-distillation teacher and negative rubrics to steer subsequent rollout generation away from recurring failure modes, addressing the lack of group-relative reward signal when all-failure and all-success rollout groups coexist in a batch. Across model backbones, RISED achieved the highest mean pass rate across environments and ranked first or second in every individual environment, according to the paper published in October 2026.

read2 min views1 publishedOct 6, 2026
RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
Image: Apple ML Research

content type paperpublished October 2026 RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

AuthorsJingtan Wang†**, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low†, Manjot Bilkhu

Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both challenges highlight limitations of relying solely on scalar rewards in multi-environment RL: they provide limited information about cross-environment relationships and no within-group reward contrast when rewards are identical. This motivates richer textual feedback, such as rubrics describing rollout behaviours, to guide learning. Beyond rubrics’ usage as reward, we repurpose rubrics to guide both online data selection and policy supervision. An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments. The resulting profiles guide the selection of data that aligns with the overall behavioural composition of the mixed-environment batch while limiting overlap with already-selected data. Available positive rubrics (describing desired behaviours) provide privileged context for an on-policy self-distillation teacher, supplying additional token-level supervision, while negative rubrics (describing undesired behaviours) guide subsequent rollout generation away from recurring failure modes. Together, these components form RISED. Across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. Rubric-based analysis of RISED can further characterize the behavioural changes accompanying these gains.

From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers August 27, 2026research area Methods and Algorithms, research area Speech and Natural Language Processing

Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during…

RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning

March 16, 2026research area Computer Vision, research area Data Science and Annotation

Dense image captioning is critical for cross-modal alignment in vision-language pretraining and text-to-image generation, but scaling expert-quality annotations is prohibitively expensive. While synthetic captioning via strong vision-language models (VLMs) is a practical alternative, supervised distillation often yields limited output diversity and weak generalization. Reinforcement learning (RL) could overcome these limitations, but its…

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @rised 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/rised-rubrics-for-ag…] indexed:0 read:2min 2026-10-06 · —