cd /news/large-language-models/compositional-reasoning-in-language-… · home topics large-language-models article
[ARTICLE · art-133315] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

A new arXiv paper (2609.19465v1) from unnamed researchers proposes a dependency-graph framework that formalizes compositional reasoning in language models into three levels of increasing complexity, and reports a consistent decomposed-to-composed asymmetry: training on decomposed skills does not reliably transfer to composed tasks, while composed-task training transfers more readily back to decomposed tasks. Using data-structure tasks with deterministic reward computation, the authors provide a theoretical explanation for the asymmetry and evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. A pilot study on real-world tool-calling benchmarks offers preliminary evidence that the decomposed-to-composed asymmetry extends to practical settings.

by read1 min views1 publishedSep 18, 2026

arXiv:2609.19465v1 Announce Type: new Abstract: Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/compositional-reason…] indexed:0 read:1min 2026-09-18 ·