Skip to content
Open access

REINFORCEMENT LEARNING WITH LOAD FORECASTING FOR SMART HOME ENERGY MANAGEMENT

Jul 2026 · Bulletin of Manash Kozybayev North Kazakhstan University · pp. 295-306 · 0 citations · 21 references

TL;DR

H-UPF is proposed, a hybrid intelligent framework for scalable sequential decision-making in heterogeneous environments under uncertainty that integrates probabilistic multi-horizon forecasting via a Temporal Fusion Transformer with continuous control via Proximal Policy Optimization, embedding predictive quantile distributions directly into the agent’s state representation.

Abstract

This paper proposes H-UPF (Hybrid Universal Policy with Forecasting), a hybrid intelligent framework for scalable sequential decision-making in heterogeneous environments under uncertainty. The architecture integrates probabilistic multi-horizon forecasting via a Temporal Fusion Transformer with continuous control via Proximal Policy Optimization, embedding predictive quantile distributions directly into the agent’s state representation. A Dynamic Adaptation Layer normalizes observations relative to instance-specific scales, enabling zero-shot policy transfer across environments with 18.5× variability in operating characteristics — without inter-agent communication or per-instance retraining. Validated on two real-world residential energy management datasets (REFIT: 20 UK households; CityLearn: 6 US buildings with real PV profiles), the framework achieves 88.4% of the theoretical optimum in zero-shot transfer, outperforming meta-learning (MAML-PPO) by 8.4 percentage points (Wilcoxon p = 0.003, Cohen’s d = 1.42). Ablation analysis identifies the adaptation layer as the dominant contributor (−16.2 p.p. upon removal), while probabilistic forecasting adds +6.8 p.p. through proactive scheduling. The learned policy is robust to reward parameter variations (≤3.2 p.p. sensitivity across 5× range) and supports practical deployment: 9.8 h one-time training, 18.4 ms inference per control step.

Read PDF

Similar papers

Conference Open access 2025

Machine Learning Predictive Models in Smart Home Energy Management: Progress and Challenges

: Home Energy Management Systems (HEMS) is becoming an essential part of the low-carbon economy and smart cities due to the global energy crisis and climate change issues. Conventional Home Energy Management Systems have significant difficulties in dealing with the complexity and heterogeneity of energy data, which hinders their practicality and wide adoption. In this comprehensive paper, the survey and review the recent progress in machine learning techniques for smart home energy management applications, highlighting key developments. The systematically categorize and thoroughly analyze the performance of various machine learning models across multiple critical tasks: energy consumption prediction, user behavior analysis, demand response optimization, and Heating, Ventilation, and Air Conditioning (HVAC) control. The findings show that deep learning methods and reinforcement learning approaches consistently achieve better prediction accuracy and enhanced adaptivity; however, persistent issues exist, such as data heterogeneity, the weak generalization ability of models across different settings, and difficulties in practical real-world deployment and implementation. To address these challenges, the strongly suggest that future research should concentrate on cross-task joint optimization strategies, multi-source data fusion techniques, and long-term adaptive model frameworks to significantly facilitate practical, intelligent energy management in homes.

Xuan-Ming Zhou · 0 citations
Open access Jul 2026

INTELLIGENT SCHEDULING OF PV–STORAGE–CHARGING INTEGRATED STATIONS VIA GTRXL-PPO WITH CURRICULUM LEARNING

To address the long-horizon sequential decision-making task, characterized by complex temporal dependencies, non-stationary dynamics, and high stochasticity in distribution-level PV–storage–charging systems, this paper develops a deep reinforcement learning framework that combines Gated Transformer‑XL (GTrXL) with Proximal Policy Optimization (PPO). Cross-segment memory captures long‑range temporal dependencies and time‑aware encodings reinforce intraday periodicity. Training adopts a five‑stage curriculum with adaptive KL control and auxiliary multi‑task heads to improve sample efficiency. A group‑normalized, potential‑based reward unifies economic performance, grid friendliness, and storage health. In simulation, the agent learns a structured six-phase daily policy and achieves a 94.4% charging completion rate and 92.8% PV utilization, reduces average daily electricity purchase cost by 15%, and keeps grid peak power within a 60 kW soft limit. Across five seeds, returns improve by 78.4% over a feed-forward PPO baseline and by 23.6% over a vanilla GTrXL-PPO, demonstrating the benefits of long-memory RL for coordinated PV–storage–charging operation. The framework enforces feasibility via continuous action mapping with ramp-rate/jerk constraints and supports millisecond-level inference. Uncertainty-aware shaping improves robustness; gains are statistically significant across five seeds via paired tests.

Song-Cheng Lu · 0 citations
Open access Aug 2026

Research on Optimal Power Grid Scheduling Based on Transfer Reinforcement Learning

To enhance power grid adaptability amid rising renewable energy integration, this paper proposes M3-PPO, a meta-reinforcement learning algorithm that enables efficient the strategy transfer and rapid adaptation across tasks with varying energy mixes. Built upon a base framework (M-PPO) that integrates PPO and MAML, M3-PPO introduces two key innovations to overcome MAML’s training instability: a Mamba-based context encoder for richer task representation in the inner loop, and a global-local momentum update mechanism for smoother meta-parameter optimization in the outer loop. Experiments on the Grid2Op platform demonstrate that M3-PPO significantly outperforms baseline algorithms in generalization and scheduling efficiency, achieving robust performance even when simulating complex energy environments. The approach is particularly suitable for integration with antenna-enabled smart grid monitoring, wireless data acquisition, and edge-computing platforms, providing an engineering-oriented solution for adaptive, real-time, and robust power grid scheduling in modern renewable-rich energy systems.

Q. Dai, X. Hu, J. Li et al. · 0 citations
Jul 2026

Sustainable Smart Farm Networks: A Decision Theory-Guided Deep Reinforcement Learning Approach

This work designs a decision theory (DT)-guided transfer learning (TL) framework that unifying cyber resilience and energy adaptability in agricultural monitoring, advancing methodological innovation with DT-guided TL for stable DRL convergence, and providing design insights for sustainable agricultural cyber-physical systems.

Dian Chen, Zelin Wan, D. Ha et al. · 0 citations
Preprint Aug 2026

DER Allocation without Load Prediction via Reinforcement Learning

A forecast-free reinforcement learning (RL) framework for DERA allocation that learns optimal policies directly from operational data, which preserves the interpretability and constraint satisfaction of DER model while adapting to stochastic demand variations through data-driven updates.

Abed AlRahman Al Makdah, Aravind Ramana, Shaofeng Zou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.