Statistical Accuracy, Economic Value and Model Instability in ETF Return Forecasting: A Comparison Across Developed and Emerging Markets
Abstract
Whether machine-learning models extract predictive signal from ETF returns across markets at different efficiency levels and whether statistical accuracy translates into trading value, remain contested. We examine this for iShares MSCI Brazil (EWZ) and iShares Core S&P 500 (IVV) from January 2010 to July 2026, training through December 2022 and testing thereafter. Random Forest, XGBoost with random search, XGBoost with Bayesian optimization, LSTM, GRU and an LSTM + XGBoost ensemble, are compared against historical mean, random walk, and AR(1) benchmarks at one-day (h = 1), five-day (h = 5) and monthly (h = 21) horizons using ten technical indicators. Every model is also evaluated against the classifier that predicts the majority class, and risk-adjusted performance is reported with bootstrap intervals. No model exceeds that trivial classifier in any combination examined. The two markets fail by distinct mechanisms: collapse onto the majority class in the developed market, and dispersed but unprofitable signals in the emerging one. Under Diebold–Mariano tests with autocorrelation-consistent variance and false-discovery control, no model is superior to the historical mean. No strategy outperforms Buy-and-Hold, and in the emerging market, no Sharpe ratio is distinguishable from zero. Where directional significance does appear, at the monthly horizon in the emerging market, it delivers no economic value. Together, these results argue for evaluating financial forecasting models simultaneously on regression metrics, economic performance, and regime stability rather than on any single criterion.