{"slug": "coskill-joint-reinforcement-learning-of-reasoning-and-meta-skill-agents-for", "title": "CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution", "summary": "Researchers introduced CoSkill, a multi-agent reinforcement learning framework that jointly trains a Reasoning Agent and a Meta-Skill Agent over a hierarchical skill library, achieving success rates of 98.4% on ALFWorld and 90.6% on WebShop, outperforming prior baselines by +3.5 and +6.2 percentage points. The framework, detailed in arXiv:2609.04865v1, enables end-to-end co-adaptation of skills and policy optimization, with code available on GitHub.", "body_md": "arXiv:2609.04865v1 Announce Type: new \nAbstract: Skill libraries improve the sample efficiency of agentic reinforcement learning (RL) by enabling large language model (LLM) agents to reuse procedural knowledge. Yet existing paradigms exhibit structural shortcomings: they either decouple skill evolution from policy optimization or instantiate meta-skills as fixed workflows. Both treat skills as passive objects to be managed, limiting the flexible evolution of skills and their co-adaptation with the reasoning agent. To address the limitations, we propose CoSkill, a unified multi-agent RL framework that recasts the static meta-skill workflow as a learnable Meta-Skill Agent and jointly trains it with a Reasoning Agent over a hierarchical skill library. By modeling the Reasoning and Meta-Skill Agents as a cooperative team sharing a single backbone, CoSkill enables end-to-end co-adaptation: the Reasoning Agent conditions its actions on a retrieved task skill and step skills selected from its child set, while its task performance guides the Meta-Skill Agent in refining those step skills. Experiments on ALFWorld and WebShop show that CoSkill substantially outperforms prior skill-based and RL baselines, achieving success rates of 98.4% and 90.6%, respectively (+3.5 and +6.2 pp). As shown in Figure 1, CoSkill achieves superior early-stage sample efficiency, asymptotic performance, and wall-clock efficiency. Our code is available at https://github.com/jinyuan-cookie/CoSkill.", "url": "https://wpnews.pro/news/coskill-joint-reinforcement-learning-of-reasoning-and-meta-skill-agents-for", "canonical_source": "https://www.machinebrief.com/news/coskill-joint-reinforcement-learning-of-reasoning-and-meta-s-qzed", "published_at": "2026-09-07 04:00:00+00:00", "updated_at": "2026-09-07 08:56:22.173772+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-agents", "ai-research"], "entities": ["CoSkill", "ALFWorld", "WebShop", "arXiv", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/coskill-joint-reinforcement-learning-of-reasoning-and-meta-skill-agents-for", "markdown": "https://wpnews.pro/news/coskill-joint-reinforcement-learning-of-reasoning-and-meta-skill-agents-for.md", "text": "https://wpnews.pro/news/coskill-joint-reinforcement-learning-of-reasoning-and-meta-skill-agents-for.txt", "jsonld": "https://wpnews.pro/news/coskill-joint-reinforcement-learning-of-reasoning-and-meta-skill-agents-for.jsonld"}}