Skip to content
Open access

Out-of-Sample Evaluation of Machine Learning for Macro-Factor-Based ETF Allocation

Jing Liu
Sep 2026 · Advances in Economics, Management and Political Sciences · 0 citations

Abstract

Whether ETF dynamic allocation can use machine learning to transform macro information into effective stock-bond signals still requires rigorous out-of-sample testing. This paper uses SPY and TLT to represent U.S. equities and long-term Treasury assets, combined with Choice prices and FRED macro data, constructing a monthly sample of 235 periods from March 2006 to September 2025, and implementing an expanding-window forecast in the last 47 periods of final holdout samples. This paper compares the direction prediction ability of Logistic Regression, Random Forest, and XGBoost, and converts the probability of unweighted Logistic Regression output into the continuous allocation weights of stocks and bonds. The results show that each model does not stably exceed the simple benchmark, and the AUC confidence intervals of the two types of Logistic Regression cover 0.5; Under the moving-block bootstrap, the probability prediction error of the unweighted model is higher than the historical prevalence benchmark. The net cumulative return of the probability-weighted strategy is 10.42%, which is higher than the static 50/50 portfolio, but lower than the buy-and-hold SPY and the same-average-weight portfolio, and the confidence interval of the strategy return difference contains zero. The research shows that the allocation value of low-frequency macro and market characteristics is limited, and strict out-of-sample testing and asset exposure control help to identify the applicable boundaries of machine learning strategies.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.