arXiv:2609.05822v1 Announce Type: new Abstract: Deploying advanced HVAC (Heating, Ventilation and Air Conditioning) controllers at scale remains difficult because performance often depends on accurate building models or per-site retuning. We propose NOMAD-RL (Neural Online Meta-Adaptation for Dynamics), a general-purpose Reinforcement Learning (RL) controller designed to transfer across heterogeneous thermal zones through a universal, non-invasive thermostat interface. The controller acts on temperature setpoints from zone measurements and forecasts, while a recurrent policy supports online adaptation under partial observability. Our main contribution is an adaptive domain randomization scheme based on physics-informed normalizing flows, which models correlated and multimodal distributions of thermal-zone parameters while maintaining physical plausibility and controllability. This produces a realistic and progressively adaptive training curriculum that improves transfer across buildings. We evaluate NOMAD-RL against a constant-setpoint PID controller, RL without domain randomization, and MPC in single- and multi-zone settings. NOMAD-RL consistently outperforms the PID and non-randomized RL baselines, and approaches the performance of a well-tuned MPC, especially in the more challenging multi-zone case. These results highlight the potential of adaptive, physics-informed domain randomization for robust and transferable HVAC control.
Generalizing HVAC Control With Domain Randomized Reinforcement Learning
Researchers proposed NOMAD-RL (Neural Online Meta-Adaptation for Dynamics), a general-purpose reinforcement learning controller that transfers across heterogeneous thermal zones through a universal, non-invasive thermostat interface, according to an arXiv paper (arXiv:2609.05822v1). The controller's main contribution is an adaptive domain randomization scheme based on physics-informed normalizing flows that models correlated and multimodal thermal-zone parameter distributions. In single- and multi-zone evaluations, NOMAD-RL consistently outperformed a constant-setpoint PID controller and RL without domain randomization, and approached the performance of a well-tuned MPC, especially in the multi-zone case.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.