Using Data-Driven Simulation Models for Deep Reinforcement Learning Based HVAC Control
Abstract
This manuscript summarizes ongoing doctoral research on control of building heating, ventilation, and air conditioning (HVAC) systems through deep reinforcement learning trained on data-driven simulators. The thesis investigates whether multivariate time-series forecasting models can act as reliable surrogates of physics-based building emulators for controller development. The work is motivated by the high modeling effort required by conventional model predictive control and by the practical impossibility of training reinforcement learning agents directly on real buildings. The proposed methodology combines synthetic data generation from established simulation frameworks, fine-tuning of time-series foundation models, and cross-platform evaluation against high-fidelity building emulators. Current progress includes a published dataset-generation study and a submitted first paper centered on zero-shot forecasting of indoor temperature and HVAC energy consumption with Tiny Time Mixers across previously unseen buildings and seasons. A later stage of the thesis will study how reinforcement learning agents can use the learned surrogates for planning and HVAC control, and whether policies trained through those models transfer back to trusted simulation environments.