{"slug": "iadd-tr-intervention-aware-dynamics-decoupling-with-targeted-regularization-for", "title": "IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning", "summary": "Researchers propose IADD-TR, a framework that decouples environment dynamics into action-intervention and action-free stages to reduce policy-induced data bias in model-based reinforcement learning. Experiments on five MuJoCo tasks show competitive returns with improved sample efficiency.", "body_md": "arXiv:2608.10634v1 Announce Type: new\nAbstract: Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making. Numerous methods have been developed to improve dynamics prediction and policy optimization for MBRL through uncertainty estimation, model regularization, and conservative value learning. However, these methods typically treat the transition model and critic as monolithic predictors, overlooking the policy-induced data bias. Consequently, action can become entangled with environmental evolution, while uneven action coverage may distort the counterfactual value estimates used for policy improvement. To address this, we propose IADD-TR, a unified framework combining Intervention-Aware Dynamics Decoupling (IADD) and Targeted Regularization (TR). IADD factorizes transitions into an action-intervention stage and an action-free natural evolution stage, using a zero-action anchor to resolve the non-uniqueness of this two-stage factorization for robust generalization. Its latent and state-aligned components are identifiable up to an invertible within-block transformation and pointwise, respectively. For policy learning, we derive TR from the efficient influence function of a replay-state policy-gradient functional. TR augments the critic with an action-density-scaled residual correction and optimizes a targeted loss, yielding doubly robust policy-gradient estimation when either the critic or the replay action density is consistently specified. Extensive experiments on five MuJoCo tasks show that IADD-TR achieves competitive returns with improved sample efficiency.", "url": "https://wpnews.pro/news/iadd-tr-intervention-aware-dynamics-decoupling-with-targeted-regularization-for", "canonical_source": "https://www.machinebrief.com/news/iadd-tr-intervention-aware-dynamics-decoupling-with-targeted-fuvq", "published_at": "2026-08-12 04:00:00+00:00", "updated_at": "2026-08-12 05:41:18.321326+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["IADD-TR", "MuJoCo"], "alternates": {"html": "https://wpnews.pro/news/iadd-tr-intervention-aware-dynamics-decoupling-with-targeted-regularization-for", "markdown": "https://wpnews.pro/news/iadd-tr-intervention-aware-dynamics-decoupling-with-targeted-regularization-for.md", "text": "https://wpnews.pro/news/iadd-tr-intervention-aware-dynamics-decoupling-with-targeted-regularization-for.txt", "jsonld": "https://wpnews.pro/news/iadd-tr-intervention-aware-dynamics-decoupling-with-targeted-regularization-for.jsonld"}}