cd /news/artificial-intelligence/de-venus-a-data-efficient-rlvr-frame… · home topics artificial-intelligence article
[ARTICLE · art-121241] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

Researchers introduced DE-Venus, a data-efficient reinforcement learning with verifiable rewards (RLVR) framework for large language models, which preserves or improves model quality using only 10% of labels or as little as 13% of relevant data across public benchmarks and three business scenarios. The framework, which unifies supervision as evolving state across data preparation and policy optimization, also reduced observed convergence steps by 63%–75% in selected business configurations, cutting annotation and training costs while maintaining scalable RL execution.

read1 min views1 publishedSep 4, 2026

arXiv:2609.03324v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, often entangling supervision logic with distributed training and hindering controlled comparison and reuse. We present DE-Venus, a unified framework for data-efficient RLVR that treats supervision as evolving state across data preparation and policy optimization. It organizes this lifecycle into three modules: Active Data Selection allocates training and annotation budgets; Weak Supervision Construction derives learning signals from unlabeled examples; and Training-Time Supervision Refinement filters or corrects unreliable supervision. DE-Venus supports seven representative methods and a data-selection pipeline by expressing method-specific decisions as offline dataset transitions or online transformations of targets, rewards, batches, and advantages while preserving verl's distributed execution contracts. Across public benchmarks and three business scenarios, separate configurations preserve or improve model quality with only 10% of labels or as little as 13% of relevant data; selected business configurations also reduce observed convergence steps by 63%--75%. DE-Venus thus reduces annotation and training costs without sacrificing scalable RL execution.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @de-venus 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/de-venus-a-data-effi…] indexed:0 read:1min 2026-09-04 ·