Jul 2026· International Journal of Applied Sciences & Development· pp. 59· 0 citations· 8 references
TL;DR
A data-driven framework that integrates labour-market microdata, industry performance metrics, and task-level automation-risk indices to forecast changes in key socioeconomic indicators is proposed, reinforcing the value of machine learning—especially Random Forest—as a robust forecasting tool for evidence-based policy in an era of rapid technological transformation.
Abstract
As industrial automation accelerates across diverse sectors, its socioeconomic repercussions remain complex and uncertain, particularly on employment, income distribution, and macroeconomic stability. This study proposes a data-driven framework that integrates labour-market microdata, industry performance metrics, and task-level automation-risk indices to forecast changes in key socioeconomic indicators. Five state-of-the-art machine learning algorithms—Random Forest, XGBoost, CatBoost, LightGBM, and TabNet—were trained and compared. These models capture complex, nonlinear interactions in socioeconomic data. Among them, Random Forest demonstrated the best predictive performance, achieving the lowest RMSE (1560.74) and highest R² (0.9690), significantly outperforming other models. Interpretable ML techniques, such as SHAP values and counterfactual simulations, are employed to identify the most influential predictors of automation-related socioeconomic change. The results offer a fine-grained, scenario-based understanding of future labour market trends, reinforcing the value of machine learning—especially Random Forest—as a robust forecasting tool for evidence-based policy in an era of rapid technological transformation.
Predictive decision analytics is gaining significance for better planning, resource allocation, and optimization of operations in business & industry. The issue of forecasting electricity prices is a relevant one as the volatility of prices directly affects procurement, production scheduling and operating costs. This study developed an artificial intelligence and machine learning framework for predictive decision analytics using electricity price forecasting as the application domain. A publicly available dataset comprising 23,304 observations and 11 variables was analysed. Historical electricity load, lagged electricity prices, day, and season were used as predictors of the current electricity price. Random Forest, XGBoost, and Gradient Boosting models were developed using an 80:20 train–test split. Model performance was evaluated using MAE, RMSE, MAPE, R², five-fold time-series cross-validation, and SHAP-based interpretation. All three models achieved comparable predictive performance. XGBoost produced the highest R² (0.8870) and lowest RMSE (927.43), while Gradient Boosting achieved the lowest MAE (575.34) and MAPE (10.44%). Time-series cross-validation reported a mean R² of 0.7823, confirming stable predictive capability under changing temporal conditions. SHAP analysis identified recent historical electricity prices, particularly P(T−1) and P(T−24), as the most influential predictors. The results prove that ensemble machine learning is a credible and explainable approach to electricity price prediction, which can be utilized for procurement planning, budgeting, production scheduling, and decision-making for business and industry. .
A. Bhargava, Vijay Agrawal, K. R. Babu et al.· Journal of Intelligent Decis...· 0 citations
The GWO-XGBoost model achieved an R² of 0.991 in Gross Domestic Product (GDP) prediction, exceeding all other machine learning models compared in this study. Thus, GWO-XGBoost provides a transparent and reliable decision-making support system for all those who develop economic policies. Developing accurate macroeconomic predictions can assist in the advancement of sustainable development. However, the non-linear relationships among many macroeconomic variables, such as industrial production, governance, and demand, make it difficult to develop accurate and understandable predictions using conventional macroeconomic models. To address the limitations of these traditional models, we propose a hybrid explanatory model, GWO-XGBoost, that integrates the Grey Wolf Optimizer (GWO) with Extreme Gradient Boosting (XGBoost). This integrated model automatically tunes its hyperparameters. Monthly US data from 1996 to 2020 was utilized in the experiment. These data include: Government Effectiveness, Consumer Demand/Retail Sales, Industrial Production, Trade Balance, and Unemployment. The results of the SHapley Additive exPlanations (SHAP) analysis show that the two most significant factors influencing economic performance were consumer demand and government effectiveness. As a result of being both highly predictive and having clear feature attributions, the proposed hybrid explanatory model will be useful for all institutions involved in developing and implementing evidence-based planning processes that align with economic growth.
Oluwaseun Racheal Ojekemi, D. Kırıkkaleli· Discover Computing· 0 citations
This study aims to reveal the interaction between environmental sustainability and health policies by examining the relationships between economic growth, carbon emissions, and healthcare expenditures within the framework of machine learning methods. The primary motivation of the research is to evaluate the relationships among economic growth, health indicators, and environmental sustainability using data-driven methods.
The analysis is based on a large-scale panel dataset covering 217 countries for the period 1960–2024, obtained from the World Bank World Development Indicators database. Advanced machine learning algorithms, including Extra Trees, Random Forest, Extreme Gradient Boosting, Histogram-based Gradient Boosting, and Light Gradient Boosting Machine, were implemented, and their prediction performances were compared. To improve prediction accuracy, a hybrid ensemble model was proposed by combining the highest-performing algorithms. For model interpretability, the relative importance levels of variables were analyzed using the Explainable Artificial Intelligence approach via Shapley Additive Explanations.
The findings indicate that the proposed hybrid ensemble model provides higher prediction accuracy compared to individual algorithms. Results from the Shapley Additive Explanations analysis show that health expenditures, economic growth, and innovation-related indicators exhibited high predictive importance within the developed machine learning models. The findings highlight the predictive associations between environmental indicators, healthcare spending, and sustainable development outcomes.
This study analyzes the complex relationships between economic growth, environmental sustainability, and health policies through a data-driven perspective. The findings provide additional evidence regarding the predictive associations among economic, environmental, and health-related indicators. The proposed framework demonstrates the usefulness of machine learning approaches for analyzing complex nonlinear patterns in large-scale country-year datasets.
Murat Tekbaş, Elif Aktepe· BMC Public Health· 0 citations
Maintenance management of stationary combustion engines in the agricultural sector remains largely manual, increasing the risk of unplanned downtime. This study developed a machine learning-based predictive model to anticipate failures within a 60-day horizon, enabling the transition from reactive to proactive maintenance. Following the CRISP-DM (Cross-Industry Standard Process for Data Mining) framework, a sliding-window feature engineering pipeline was built from 2250 historical records spanning 59 engines. Four ensemble learners (Random Forest, LightGBM, XGBoost, and CatBoost) were then compared under two complementary protocols: a strict 60/40 chronological split simulating deployment, and a stratified leave-engine-group-out cross-validation withholding entire engines from training. Nonparametric testing (DeLong test and engine-level cluster bootstrap) showed that the four learners are statistically equivalent, whereas the feature engineering layer contributes a large, significant discrimination gain (ΔAUC ≈+0.06, p≈10−27) on engines unseen during training. Random Forest, selected as the final model, achieved an AUC of 0.90 with 84.2% recall under the deployment protocol and 0.96 with 90.7% recall under engine-grouped validation. Temporal extrapolation, rather than cross-engine generalization, appeared to be the primary challenge, indicating that rigorously engineered degradation features, more than the choice of ensemble algorithm, drive predictive performance in agricultural maintenance planning.
I. Jaramillo, Walter Orozco-Iguasnia, Rubén Patricio Alcocer Quinteros et al.· Algorithms· 0 citations
Data-driven LGT can help identify meaning and reshape the provision of public services, and facilitate expansion in new global and digital business sectors. Hence, the value of machine learning (ML)-based predictive intelligence has increased in the context of providing actionable predictions based on real-life data and policy recommendations. In this paper, a comprehensive machine learning framework is proposed to support the decision-making and policy-making processes in a more advanced manner. The framework includes several prediction algorithms such as Linear Regression, Decision Trees, Random Forest and Artificial Neural Networks to make predictions about the important information provided by the government. The models are demonstrated and analyzed using real-world data from open government portals across different scenarios, like prediction of budget, hospital demand forecasting, as well as emergency source allocation. They are tested for their predictive efficiency with traditional metrics such as the Root Mean Square Error (RMSE), the Mean Absolute Error (MAE) and the Coefficient of Determination (R²). A comparison of the experimental results indicates that Artificial Neural Networks have significant predictive performance, especially for complex and high-dimensional data, whereas Random Forest has a high predictive performance and a better interpretability. Linear Regression and Decision Trees are more useful in terms of transparency and scalability, making them appropriate models for selection of public policies in government systems. They are not a good environment for implementation everywhere, however, including universities, NGOs, or national/international government institutions with many requirements for their implementation. This encompasses application of explainable artificial intelligence (XAI) methods to tackle transparency and trust concerns, in addition to designing more local and adaptive models to enhance generalizability to various administrative contexts.
The global energy transition and ongoing electricity market reforms require power grid enterprises to balance reliable power supply with increasingly stringent transmission and distribution tariff regulations. Traditional budgeting methods based on historical extrapolation often fail to reflect the physical basis of asset operations, whereas data-driven machine learning models achieve high predictive accuracy but lack the transparency required for regulatory cost verification. To address the trade-off between forecasting accuracy and interpretability, this study proposes a hybrid cost-forecasting model based on cost quotas. The framework uses standardized operating quotas as the physical budgeting baseline and incorporates a dynamic mechanism for quota evolution driven by macroeconomic conditions and technological progress. Extreme Gradient Boosting (XGBoost) is employed to capture nonlinear residuals beyond the quota-based estimates, while SHAP (Shapley Additive exPlanations) is used to interpret the contribution of key cost drivers. The model was evaluated using 16 years of anonymized operational data from a provincial power grid in China. It achieved a mean absolute percentage error (MAPE) of 2.34%, reducing forecasting errors by 61.8%, 46.6%, and 34.1% compared with SARIMAX, standalone XGBoost, and Attention-LSTM models, respectively. The proposed framework integrates engineering cost-quota principles with explainable artificial intelligence, providing both accurate long-term cost forecasts and a transparent decision-support tool for regulatory permitted-cost verification.
Xiaohui Wang, Tong Li, Yanchao Lu et al.· Journal of Visualized Experi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.