Skip to content
Preprint

Forecasting and Explaining the Phillips Curve: A SHAP-Based Comparison of Machine Learning and Traditional Time-Series Models for Canadian Unemployment and Inflation

Jul 2026 · 0 citations
Mathematics

TL;DR

Analysis of the top-performing XGBoost model indicates that lagged inflation is more influential than unemployment, which only becomes significantly impactful during the pandemic tail, and clarify when machine learning methods can surpass traditional benchmarks.

Abstract

This study evaluates the out-of-sample forecasting ability of six model types: ARIMA, VAR, Random Forest, XGBoost, LSTM, and GRU, for monthly Canadian inflation from January 2012 to April 2026 (n = 172). The evaluation employs expanding-window walk-forward validation across 1-, 3-, 6-, and 12-month horizons. Results reveal a horizon-dependent shift: ARIMA significantly outperforms all machine learning and deep learning models at the one-month horizon (Diebold-Mariano p<0.05). However, Random Forest and XGBoost become notably superior at six and twelve months, reducing RMSE by approximately 30-75 percent compared to ARIMA and VAR. LSTM and GRU perform well only at the shortest horizon, likely due to overfitting given the limited data. Analyzing four macroeconomic sub-periods shows that no single model consistently dominates. SHAP analysis of the top-performing XGBoost model indicates that lagged inflation is more influential than unemployment, which only becomes significantly impactful during the pandemic tail. The findings clarify when machine learning methods can surpass traditional benchmarks.

View source

Similar papers

Open access Aug 2026

Forecasting Multivariate Time Series: A Comparison of Machine Learning, Statistical and Deep Learning Models

The findings demonstrate that rigorous leakage-free validation is essential for reliable forecasting research and that, for monthly Robusta coffee prices, increased model complexity does not necessarily yield superior predictive performance.

Dler H Kadir, D. Khalil, Azhin M. Khudhur · 0 citations
2025

Machine Learning-Based Forecasting of Multivariate Time Series: Evidence from Random Forest and Extreme Gradient Boosting for VAR Model

AbstractConventional Vector Autoregressive (VAR) models are widely applied for multivariate time series analysis, their performance deteriorates in high-dimensional settings due to inefficient parameter estimation, unstable forecasts, and difficulties in interpreting temporal dependencies. This study conducts a comparative study on conventional Vector Autoregressive (VAR) models, Multivariate Random Forest for VAR (MRF-VAR) models and Multivariate Extreme Gradient Boosting for VAR (MXGB-VAR) models, validated using simulated and real-life dataset for Nigerian financial time series. Augmented Dickey–Fuller (ADF) test was adopted to test the stationarity of the data. Forecast accuracy across short-term and long-term horizons for the models were measured using Mean Absolute Deviation (MAD) and Root Mean Square Deviation (RMSD). Results for simulated data show that the conventional VAR models achieved the best short-term forecast performance (MAD = 1.642, RMSD = 2.016), while the MRF-VAR models performed best in long-term forecasting (MAD = 0.947, RMSD = 1.197). In the case of the real-life dataset for Nigerian financial time series, the MRF-VAR model performed well in short-term forecasts (MAD = 108.84, RMSD = 149.53), whereas MXGB-VAR model provided better results in long-term forecasts (RMSD = 730.57). Policymakers and financial analysts should be encouraged to apply machine learning approaches to VAR models in macroeconomic and financial forecasting to improve decision-making.

N. Isah, S. Doguwa · 0 citations
Open access Aug 2026

Comparing machine learning and classical methods in forecasting Türkiye’s macroeconomic performance

This study forecasts Türkiye’s medium-term macroeconomic performance through a composite index based on growth, unemployment, inflation, the budget balance, and the current account balance. Quarterly data for 2006Q1–2026Q1 are used to compare artificial neural networks, ARIMA/SARIMA models, and ordinary least squares regression. Model performance is evaluated through rolling-origin validation over 40 out-of-sample periods using MAE, MAPE, and RMSE. The statistical significance of forecast error differences is examined with the Diebold–Mariano test, while the sensitivity of the neural network is assessed across 180 hyperparameter configurations. The results show that the neural network produces the lowest error in the baseline specification, although its advantage is not statistically significant at the 5 percent level. When seasonal information and lagged component values are included, OLS yields the lowest forecast error. Conditional forecasts from the preferred OLS specification place the index between 91.66 and 94.56 during 2026Q2–2028Q1. The projected path remains broadly stable, with seasonal fluctuations but no pronounced upward or downward trend. Overall, the findings indicate that forecast performance depends on model specification and that claims of machine learning superiority should be evaluated cautiously.

Emrah Kıratoğlu · 0 citations
Preprint Aug 2026

Forecasting in the Fog: Real-Time versus Revised-Data Evidence on Machine Learning's Edge over the Phillips Curve

ML inflation forecasts are almost universally trained on fully revised data, even though real-time forecasters never have such data, and reported feature importances are typically computed in-sample, conflating predictive relevance with retrospective fit. This paper asks whether the ML advantage over the Phillips curve documented in Agyekum (2026) survives when models are trained and evaluated on real-time (ALFRED) vintages rather than revised series, and whether SHAP feature-importance rankings are an artifact of in-sample estimation. Using 2000-2026 U.S. data on unemployment, CPI and PCE inflation, payrolls, real GDP, and the 10-year-2-year Treasury spread, vintage-consistent panels are built for four traditional models (random walk, AR(1), Phillips curve, ADL-OLS) and four ML models (Random Forest, Gradient Boosting, Elastic Net, SVR), re-estimated recursively at 3-, 6-, and 12-month horizons (208, 206, 204 forecasts). Real-time/revised accuracy differences are small and, apart from one exception at 6 months (Gradient Boosting vs. Phillips curve, DM = -1.671, p = 0.097), indistinguishable under Diebold-Mariano tests; Gradient Boosting alone shows consistent positive skill at longer horizons. The random walk remains a strong short-horizon benchmark, consistent with the puzzle in Agyekum et al. (2026) for exchange rates. Using walk-forward, out-of-sample SHAP, a Random Forest on revised data assigns dominant importance to PCE inflation (mean |SHAP| = 0.778, rank 1 of 9), while on real-time data it assigns PCE negligible importance (0.039, rank 6), relying instead on current CPI (0.834 vs. 0.223). This twenty-fold swing, larger than the in-sample estimate, is invisible to point-forecast metrics and shows the model's PCE reliance is substantially a hindsight artifact. An RSI summarizes the accuracy gap by model and horizon, with implications for auditing ML inflation forecasts.

Louis Agyekum, Obed Obese · 0 citations
Open access Sep 2026

Out-of-Sample Evaluation of Machine Learning for Macro-Factor-Based ETF Allocation

Whether ETF dynamic allocation can use machine learning to transform macro information into effective stock-bond signals still requires rigorous out-of-sample testing. This paper uses SPY and TLT to represent U.S. equities and long-term Treasury assets, combined with Choice prices and FRED macro data, constructing a monthly sample of 235 periods from March 2006 to September 2025, and implementing an expanding-window forecast in the last 47 periods of final holdout samples. This paper compares the direction prediction ability of Logistic Regression, Random Forest, and XGBoost, and converts the probability of unweighted Logistic Regression output into the continuous allocation weights of stocks and bonds. The results show that each model does not stably exceed the simple benchmark, and the AUC confidence intervals of the two types of Logistic Regression cover 0.5; Under the moving-block bootstrap, the probability prediction error of the unweighted model is higher than the historical prevalence benchmark. The net cumulative return of the probability-weighted strategy is 10.42%, which is higher than the static 50/50 portfolio, but lower than the buy-and-hold SPY and the same-average-weight portfolio, and the confidence interval of the strategy return difference contains zero. The research shows that the allocation value of low-frequency macro and market characteristics is limited, and strict out-of-sample testing and asset exposure control help to identify the applicable boundaries of machine learning strategies.

Jing Liu · 0 citations
Open access Aug 2026

A COMPARATIVE ANALYSIS OF FORECASTING ACCURACY BETWEEN MACHINE LEARNING MODELS AND OLS REGRESSION: EMPIRICAL EVIDENCE FROM THE VIETNAMESE STOCK MARKET

This research investigates and compares the predictive performance of stock return forecasting between the Ordinary Least Squares (OLS) regression model and machine learning approaches in the context of the volatile Vietnamese stock market. Using a panel dataset combined with time-series data of listed firms on the Vietnamese stock exchange from 2015 to 2024, the study contrasts the OLS model with three advanced machine learning algorithms, including Artificial Neural Networks (ANN), Random Forest, and XGBoost. Predictive performance is primarily assessed using Root Mean Squared Error (RMSE), alongside Mean Absolute Error (MAE) and out-of-sample R-squared (R²OOS) as robustness measures. The empirical results demonstrate that machine learning models significantly outperform OLS in capturing complex nonlinear relationships in stock returns. Among them, the ANN model achieves the lowest RMSE, indicating the highest predictive accuracy, and generates superior long–short portfolio returns compared to the other models. Furthermore, the Diebold–Mariano test confirms that the differences in predictive accuracy between machine learning models and OLS are statistically significant. Although OLS retains advantages in terms of simplicity and interpretability, machine learning models exhibit clear superiority in predictive performance and quantitative risk management. This study provides important empirical evidence from an emerging market such as Vietnam and offers practical implications for investors and policymakers in optimizing asset allocation decisions.

Phat Ly Huynh Ngo, T. Pham, B. Lệ · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.