cd /news/machine-learning/dace-dt-data-centric-offline-multi-t… · home › topics › machine-learning › article
[ARTICLE · art-148062] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

DaCe-DT: Data-Centric Offline Multi-Task Reinforcement Learning via Adaptive Prompts and Trajectory Correction for Heterogeneous Tasks

A new arXiv paper (arXiv:2610.11085v1) introduces DaCe-DT, an offline multi-task reinforcement learning framework that combines length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC) to address three data-level bottlenecks: ineffective prompt-length use, semantic irrelevance of randomly sampled prompt segments, and misleading supervision from fragmented trajectories. On Meta-World benchmarks, DaCe-DT outperformed state-of-the-art methods by an average of 11.73% on optimal datasets and 13.34% on suboptimal datasets.

by read1 min views1 publishedOct 9, 2026

arXiv:2610.11085v1 Announce Type: new Abstract: Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data. However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance. This paper, from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:(i) ineffective utilization of prompts length under diverse task complexities, and (ii) semantic irrelevance of randomly sampled prompt segments, (iii) misleading supervision induced by fragmented and discontinuous trajectories. To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality, featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC). Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks. Experimental results on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods, achieving an average improvement of 11.73% on optimal datasets and an improvement of 13.34% on suboptimal datasets, demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.

── more in #machine-learning 4 stories · sorted by recency
── more on @dace-dt 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dace-dt-data-centric…] indexed:0 read:1min 2026-10-09 · —