cd /news/machine-learning/learning-hierarchical-skill-policies… · home topics machine-learning article
[ARTICLE · art-105424] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Learning Hierarchical Skill Policies with Offline Quality-Diversity Reinforcement Learning

Researchers introduced QDOS (Quality-Diversity Offline Skill learning), a unified pipeline for offline-to-online reinforcement learning that uses an Advantage-Weighted Quality-Diversity pretraining objective to extract diverse and high-value skills from pre-collected datasets. In experiments, QDOS significantly outperformed strong baselines in structured manipulation and unstructured locomotion tasks, accelerating exploration and improving final returns in sparse-reward domains.

read1 min views3 publishedAug 21, 2026

arXiv:2608.19684v1 Announce Type: new Abstract: Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL. One promising approach to achieve this goal is to employ a two-stage strategy: In the first stage, diverse skills are extracted as a low-level policy from a given dataset, and a high-level policy is trained to solve a specific task in the second stage. Typically, extraction of the low-level policy is performed based on unsupervised learning such as trajectory VAE. However, a limitation of this approach is that the quality of the low-level policy highly depends on the quality of the dataset. To address this issue, we introduce QDOS (Quality-Diversity Offline Skill learning), a unified pipeline for robust offline-to-online learning. Our approach incorporates an Advantage-Weighted Quality-Diversity pretraining objective, which weights the skill extraction and diversity objectives by the estimated advantage of each trajectory segment. This approach allows the model to extract diverse and high-value skills. By providing robust and task-relevant skill representations, QDOS significantly improves the quality of the embedded skill space used by the low-level policy. We further integrate this with a dual dataset reuse strategy, where offline data is used both for skill pretraining and for populating the online replay buffer via pseudo-labeling. Experiments demonstrate that QDOS significantly outperforms strong baselines in structured manipulation tasks and unstructured locomotion tasks, confirming its ability to accelerate exploration and improve final returns in challenging sparse-reward domains.

── more in #machine-learning 4 stories · sorted by recency
── more on @qdos 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/learning-hierarchica…] indexed:0 read:1min 2026-08-21 ·