{"slug": "generalizing-hvac-control-with-domain-randomized-reinforcement-learning", "title": "Generalizing HVAC Control With Domain Randomized Reinforcement Learning", "summary": "Researchers proposed NOMAD-RL (Neural Online Meta-Adaptation for Dynamics), a general-purpose reinforcement learning controller that transfers across heterogeneous thermal zones through a universal, non-invasive thermostat interface, according to an arXiv paper (arXiv:2609.05822v1). The controller's main contribution is an adaptive domain randomization scheme based on physics-informed normalizing flows that models correlated and multimodal thermal-zone parameter distributions. In single- and multi-zone evaluations, NOMAD-RL consistently outperformed a constant-setpoint PID controller and RL without domain randomization, and approached the performance of a well-tuned MPC, especially in the multi-zone case.", "body_md": "arXiv:2609.05822v1 Announce Type: new \nAbstract: Deploying advanced HVAC (Heating, Ventilation and Air Conditioning) controllers at scale remains difficult because performance often depends on accurate building models or per-site retuning. We propose NOMAD-RL (Neural Online Meta-Adaptation for Dynamics), a general-purpose Reinforcement Learning (RL) controller designed to transfer across heterogeneous thermal zones through a universal, non-invasive thermostat interface. The controller acts on temperature setpoints from zone measurements and forecasts, while a recurrent policy supports online adaptation under partial observability.\n  Our main contribution is an adaptive domain randomization scheme based on physics-informed normalizing flows, which models correlated and multimodal distributions of thermal-zone parameters while maintaining physical plausibility and controllability. This produces a realistic and progressively adaptive training curriculum that improves transfer across buildings. We evaluate NOMAD-RL against a constant-setpoint PID controller, RL without domain randomization, and MPC in single- and multi-zone settings. NOMAD-RL consistently outperforms the PID and non-randomized RL baselines, and approaches the performance of a well-tuned MPC, especially in the more challenging multi-zone case. These results highlight the potential of adaptive, physics-informed domain randomization for robust and transferable HVAC control.", "url": "https://wpnews.pro/news/generalizing-hvac-control-with-domain-randomized-reinforcement-learning", "canonical_source": "https://arxiv.org/abs/2609.05822", "published_at": "2026-09-10 04:00:00+00:00", "updated_at": "2026-09-10 04:23:13.332283+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "ai-safety"], "entities": ["NOMAD-RL", "arXiv", "PID controller", "MPC"], "alternates": {"html": "https://wpnews.pro/news/generalizing-hvac-control-with-domain-randomized-reinforcement-learning", "markdown": "https://wpnews.pro/news/generalizing-hvac-control-with-domain-randomized-reinforcement-learning.md", "text": "https://wpnews.pro/news/generalizing-hvac-control-with-domain-randomized-reinforcement-learning.txt", "jsonld": "https://wpnews.pro/news/generalizing-hvac-control-with-domain-randomized-reinforcement-learning.jsonld"}}