Skip to content
Review

Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical Study

Jul 2026 · arXiv.org · Vol abs/2607.14024 · 0 citations · 22 references
Computer Science

TL;DR

Cluster-based Sequential Feature Selection (CSFS), a novel, model-agnostic, clustering-based wrapper method for automatic, efficient, and reliable feature selection in renewable energy prediction pipelines, is proposed and empirically evaluated on both use cases.

Abstract

With rising global energy demand and growing awareness of climate change and its impacts, the share of renewable energies in the global energy mix continues to grow. Unlike conventional power generation, the output of renewable energy sources cannot be controlled as consistently due to their dependence on environmental conditions. Therefore, reliable prediction of current and future energy production is essential. In this paper, we report findings from two structured literature reviews on real-world renewable energy prediction tasks: wind turbine power curve modeling and photovoltaic power prediction. For the former, we conducted a comprehensive literature review ourselves, while for the latter, we synthesize the key findings regarding frequently selected input features based on an existing survey. Across both domains, our analysis reveals that despite the large number of available monitoring and environmental variables, only limited or unsystematic methods for feature selection exist. To address this gap, we propose Cluster-based Sequential Feature Selection (CSFS), a novel, model-agnostic, clustering-based wrapper method for automatic, efficient, and reliable feature selection in renewable energy prediction pipelines. To support reproducibility and reuse, we provide an open-source implementation of CSFS on GitHub. We empirically evaluate the proposed approach on both use cases and compare it with established feature selection techniques such as wrapper-based sequential feature selection (SFS), filter-based methods, and Random Forest's embedded feature importance. The results show that the wrapper-based methods overall provide better-performing selections of features. CSFS achieves a predictive performance comparable to SFS while reducing computational cost by an average of 21%.

View source

Similar papers

#explainable ai Review Open access Aug 2026

AI-driven forecasting for efficient integration of renewable energy systems

Emerging research directions, such as explainable AI, federated learning, digital twins, edge intelligence, and physics-informed machine learning, are identified as promising strategies for developing resilient, intelligent, and sustainable future power grids.

Olatunde Ibiyinka, Tolu Omotoso, N. Ekekwe · 0 citations
Review Open access Aug 2026

Application and Comparative Study of Time Series Analysis Algorithms in New Energy Forecasting

The intermittent and volatile characteristics of new energy generation, together with the increasing demand for stable power supply in intelligent industrial systems, make accurate forecasting a critical issue for grid dispatch and electromagnetic energy management. This study systematically reviews the technological evolution of time series analysis methods for wind and photovoltaic (PV) power forecasting and establishes a comparative framework covering classical statistical models, intelligent learning algorithms, and hybrid modeling strategies. Based on two years of operational data collected from an actual wind farm and PV station in East China, the forecasting performance of ARIMA, exponential smoothing, Support Vector Regression (SVR), Long Short-Term Memory (LSTM) networks, and Transformer architectures is comprehensively evaluated, while hybrid approaches based on Empirical Mode Decomposition (EMD) are further investigated. The results demonstrate that model selection should jointly consider forecasting horizon, data characteristics, and computational constraints. Classical statistical methods remain robust under stable operating conditions but are less effective in capturing extreme fluctuations, whereas deep learning approaches exhibit superior capability in modeling long-range temporal dependencies despite reduced interpretability. Decomposition-based hybrid strategies achieve a more balanced performance across diverse scenarios and show enhanced robustness under extreme weather conditions. The study further proposes a structured model selection guideline by matching data characteristics with operational requirements, providing theoretical support for forecasting system design in renewable-energy-driven power networks and offering useful references for electromagnetic energy utilization and intelligent industrial applications.

M. S. Song, C. Yang, Z. Heng et al. · 0 citations
Review Open access Jul 2026

A review of machine learning and deep learning models, features, and applications for solar PV forecasting

Solar energy has become an important source of renewable energy towards supporting the increased electricity demand in the world and minimizing reliance on fossil energy. Nevertheless, solar irradiance is intermittent and unreliable, which necessitates the precise prediction of solar energy to promote effective integration and energy management into the grid. This review bridges a gap in the research to unify the different machine learning (ML) and deep learning (DL) methods to predict solar power and solar resource, with the need to have a comparative evaluation of the performance, interpretability, and adaptability of the methods. This paper presents a systematic review of the recent developments in ML and DL based models applied in the analysis of solar power and solar resources forecasting, such as standalone, hybrid, and ensemble models. The research topic is to determine the appropriate parameters of input, feature selection techniques, and forecasting horizons that improve the accuracy and precision of the model. The methodology that was adopted is a comprehensive comparative analysis of the previous research with an emphasis on the strengths, weaknesses, and possibilities of the various predictive models. The findings prove that hybrid and ensemble models are invariably more successful compared to traditional models, in terms of accuracy and reliability. The review shows that the synergy of AI-based solutions can determine significant improvements in terms of accuracy of forecasting, energy scheduling, and grid stability. Taking all of this into consideration, the provided review fits into the smart prediction model evolution for effective and environmentally friendly use of solar energy. The key contributions of this paper and the main highlights are as follows: Introduces the general classification of ML and DL methods applied on PV power and solar resource prediction. Compares single, hybrid, and collaborative forecasting destructions at diverse timescales. Discusses how feature selection and meteorological parameters can be used to improve the accuracy of the model. Recognizes interpretability, data availability, as well as generalization challenges in AI-based forecasting. Recommends future research initiatives, such as transfer learning, probabilistic forecasting, and explainable AI usages. Proposed a graphical taxonomy to have a better conceptual knowledge and viable model choice. Provides an updated synthesis (2020–2025) of literature to ensure contemporary relevance to solar PV forecasting research. Introduces the general classification of ML and DL methods applied on PV power and solar resource prediction. Compares single, hybrid, and collaborative forecasting destructions at diverse timescales. Discusses how feature selection and meteorological parameters can be used to improve the accuracy of the model. Recognizes interpretability, data availability, as well as generalization challenges in AI-based forecasting. Recommends future research initiatives, such as transfer learning, probabilistic forecasting, and explainable AI usages. Proposed a graphical taxonomy to have a better conceptual knowledge and viable model choice. Provides an updated synthesis (2020–2025) of literature to ensure contemporary relevance to solar PV forecasting research.

Dheeraj Sharma, C. Kumar, Manish Kumar · 0 citations
Conference Aug 2026

Following rigorous accuracy and research on constructing scenario sets for combined wind and solar power output based on kernel density estimation and copula functions

The inherent uncertainties of renewable energy sources—specifically their intermittent and volatile nature—create significant obstacles for power system planning. With the growing integration of renewables into the grid, mapping the precise interdependencies among wind generation, solar photovoltaic (PV) yield, and power demand has become crucial for designing and managing resilient power networks. To address the challenge of simulating representative scenarios for correlated wind and solar outputs, this study initially employs non-parametric kernel density estimation (KDE) to model extensive empirical datasets. Following rigorous accuracy and goodness-of-fit validation, specific KDE formulas for both wind and solar resources are derived. Subsequently, the research constructs various Copula-based joint probability models to represent the combined power generation of wind and solar farms. The performance of these models is comparatively assessed using Maximum Likelihood Estimation (MLE) alongside Akaike and Bayesian Information Criteria (AIC/BIC). reduce nce, the most suitable Copula function is identified to capture the joint probabilistic behavior of wind and PV systems. Ultimately, this optimal Copula model drives the generation of annual power output profiles for wind and solar energy. Computational experiments and validation procedures confirm that the simulated yearly scenarios accurately preserve the underlying correlation structures of the original data.

Lijian Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.