Skip to content
Conference

Reinforcement Learning-Based Building Energy Optimization for Flexibility Markets

Jul 2026 · 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET) · pp. 1-4 · 0 citations · 16 references

Abstract

The rapid growth of energy markets and demand-side response programs has created a significant need for intelligent building-level control strategies to capture high volatility energy consumption in response to price signals and grid conditions. This paper presents a reinforcement learning (RL)-based building energy management framework that models commercial buildings as active, grid-interactive assets capable of providing real-time flexibility while maintaining occupant comfort. The proposed approach integrates historical and real-time data from IoT sensors, HVAC systems, and weather forecasts to build an adaptive environment for RL agents. The RL model is applied to learn optimal control policies that minimize operational energy cost in response to flexibility markets through load shifting, peak shaving, and short-term demand response actions. The framework also incorporates a forecasting module to predict 30-minute interval energy consumption using deep learning, enabling proactive decision-making under uncertainty. Results from simulation experiments demonstrate that the RL agents achieve significant cost savings compared to rule-based control strategies and offer a reliable, automated control to unlock underlying flexibility within building systems. The paper discusses problem formulation, algorithmic development, simulation workflows, comparative metrics, and practical deployment considerations for Saudi Arabia's smart city initiatives.

View source

Similar papers

Deep reinforcement learning-based optimal energy management for microgrids

Microgrids play a critical role in enhancing the flexibility, reliability, and sustainability of modern power systems by integrating distributed energy resources, energy storage systems, and controllable loads. However, the inherent uncertainty of renewable generation and the stochastic nature of load demand pose significant challenges to optimal energy management. To address these issues, this paper proposes a deep reinforcement learning (DRL)-based optimal energy management framework for microgrids. The problem is formulated as a Markov decision process, where the system state captures renewable generation, load demand, and storage status, while the control actions determine power dispatch and energy storage operation. A deep reinforcement learning model is developed to learn optimal control policies through continuous interaction with the environment, enabling adaptive decision-making under dynamic and uncertain conditions. To improve learning efficiency and policy stability, state normalization and reward shaping strategies are incorporated. Furthermore, a constrained optimization mechanism is introduced to ensure operational safety and economic feasibility. Experimental results on benchmark microgrid scenarios demonstrate that the proposed method outperforms conventional rule-based strategies and model-based optimization approaches in terms of operational cost reduction, energy utilization efficiency, and robustness under uncertainty. The results indicate that the proposed DRL-based framework provides an effective and scalable solution for intelligent microgrid energy management.

Li Chen, Hong-Qiao Li, Zhen-Xing Chen et al. · 0 citations
Open access Jul 2026

AI-driven optimization: revolutionizing energy efficiency in modern buildings.

This work establishes a novel continuous-time dynamic-policy learning paradigm that integrates predictive modeling with real-time adaptive control, advancing data-driven intelligent building operation toward sustainable and autonomous energy management.

Hamoud H. Alshammari · 0 citations
Open access 2026

Deep Reinforcement Learning-Based Adaptive Switching for Risk-Cost-Optimized Renewable Smart Grids

The increasing penetration of weather-driven renewable energy sources in smart grids introduces operational instability, harmonic distortion, and elevated switching costs due to the limitations of rule-based and deterministic control strategies. This study proposes a deep reinforcement learning-based adaptive switching framework to enhance renewable utilization while minimizing operational risk and economic cost. A simulation-derived dataset incorporating renewable generation, load demand, total harmonic distortion, voltage deviation, frequency variation, and risk–cost indices was generated from a risk–cost optimized smart grid model and implemented in Google Colab. The switching problem was formulated as a Markov decision process with a state space composed of power quality and economic variables, and a discrete action space representing operational modes. A Deep Q-Network agent was trained over 24-hour episodes to learn optimal switching policies. Comparative evaluation against conventional, rule-based, and analytical risk–cost optimization strategies demonstrated up to 14% reduction in total harmonic distortion, 22% reduction in switching frequency, 17% reduction in operational cost, and 19% improvement in risk mitigation, while increasing renewable penetration by 11%. The proposed framework provides a scalable and intelligent solution for industrial smart grid applications.

M. Meyyappan, P. Avirajamanjula, P. Marimuthu et al. · 0 citations
Open access 2021

Reinforcement Learning for Smart Grid Energy Optimization

The study concludes that reinforcement learning will be a highly robust and versatile solution to next-generation optimization of cyber grids with regard to data-based and autonomous, data-based grid management systems.

Fatima Zahra El Idrissi · 1 citation
Open access Aug 2026

Hierarchical Reinforcement Learning for Integrated Energy System Scheduling Based on Large Language Model Forecasting

The uncertainties in source-side renewable power supply and user-side multi-energy demand pose significant challenges to coordinated scheduling in an electricity–heat–hydrogen integrated energy system (EHH-IES). A hierarchical scheduling approach for EHH-IES is introduced, with source–load forecasts serving as its basis. Traditional forecasting methods heavily rely on large amounts of training samples. To address the forecasting challenge in data-scarce scenarios, a frozen large language model assisted by variational mode decomposition is developed for joint source–load forecasting. During the scheduling process, conventional single-level reinforcement learning strategies are not sufficiently effective in dealing with the high-dimensional hybrid action space while satisfying the intricate operating constraints of the EHH-IES. Therefore, an imitation-learning-based hierarchical proximal policy optimization strategy is developed to decompose the scheduling task into system-level energy coordination and device-level action execution. The experimental evaluation shows that the proposed forecasting approach delivers improved forecasting accuracy in data-scarce scenarios, reducing the RMSE of photovoltaic power, wind power, electric load, heat load, and hydrogen load forecasting by 8.65%, 23.27%, 34.99%, 24.60%, and 39.38%, respectively, compared with the strongest baselines. The proposed scheduling strategy achieves the fastest convergence compared with the three benchmark methods while reducing the total operating cost by 21.3%, 7.1%, and 2.8%, respectively.

Ruo-Xu Zhao, Xuan Tan, Hui Wei et al. · 0 citations
Preprint Aug 2026

Safe Deep Reinforcement Learning for Energy-Efficient HVAC Control in Multi-Zone Residential Buildings

HVAC systems represent a major share of building energy consumption. Traditional control strategies are limited in coordinating energy-comfort tradeoffs across multiple zones simultaneously. Reinforcement learning (RL) offers adaptive, data-driven control that optimizes performance over time. However, deploying learned neural network controllers in safety-critical building systems remains challenging due to lack of formal safety guarantees. We propose a safety-certified deep RL framework for multi-zone residential HVAC control. Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) agents are trained in an EnergyPlus/Sinergym simulation to minimize energy consumption while maintaining thermal comfort. Post-training safety certification is performed on the PPO policy using Lipschitz-based forward invariance analysis, building on existing tools for the computation of Lipschitz constants for neural networks, to guarantee constraint satisfaction. Both agents are evaluated over an annual simulation cycle in an eight-zone variable refrigerant flow (VRF) testbed. The PPO agent achieves 67\% comfort violation reduction compared to rule-based control, while the SAC agent achieves 27.6\% energy savings. The PPO policy satisfies formal safety certification with a margin of $2.003^\circ$C. These results demonstrate the feasibility of combining reinforcement learning with post-training safety verification for multi-zone building control.

Oussama Ziadi, A. Rochd, S. I. Kaitouni et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.