Skip to content
Open access

Forecasting Multivariate Time Series: A Comparison of Machine Learning, Statistical and Deep Learning Models

Aug 2026 · Forecasting · 0 citations · 42 references

TL;DR

The findings demonstrate that rigorous leakage-free validation is essential for reliable forecasting research and that, for monthly Robusta coffee prices, increased model complexity does not necessarily yield superior predictive performance.

Abstract

This study develops a rigorous, leakage-free forecasting framework for monthly Robusta coffee prices using historical observations from January 1975 to December 2025. A comprehensive set of explanatory variables is constructed from lagged coffee prices, moving averages, logarithmic returns, rolling volatility, and exogenous variables such as the Oceanic Niño Index (ONI), the U.S. Dollar Index, and Brent crude oil prices. To ensure methodological fairness, all predictors are generated exclusively from information available at the forecast origin, and all competing models are evaluated under a unified expanding-window walk-forward validation framework. Seven forecasting models are compared: Naïve, Exponential Smoothing (ETS), ARIMA, ARIMAX, Extreme Gradient Boosting (XGBoost), Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU). Forecasting performance is evaluated using R2, RMSE, MAE, and MAPE, while Taylor diagrams and the Diebold–Mariano test are employed to assess model agreement and differences in predictive accuracy. The results show that XGBoost achieves the highest forecasting accuracy (R2 = 0.956, RMSE = 0.264), followed closely by the Naïve (R2 = 0.954, RMSE = 0.271) and ARIMA (R2 = 0.954, RMSE = 0.270) benchmarks, whereas ARIMAX and ETS provide comparable performance and the deep learning models (LSTM and GRU) produce substantially larger prediction errors. Feature importance analysis further indicates that the first lag of coffee price is the dominant predictor, accounting for approximately 94% of the predictive gain in XGBoost. Overall, the findings demonstrate that rigorous leakage-free validation is essential for reliable forecasting research and that, for monthly Robusta coffee prices, increased model complexity does not necessarily yield superior predictive performance.

Read PDF

Similar papers

Open access Aug 2026

Comparative Analysis of Statistical, Machine Learning, and Deep Learning Models for USD/IDR Prediction Using Macroeconomic Indicators and Explainable Artificial Intelligence

The USD/IDR exchange rate is a key daily barometer of Indonesia's economic health. Accurate forecasting is vital for trade, inflation, and monetary stability. However, its volatile and nonlinear dynamics pose challenges. While research has applied statistical models, machine learning, and deep learning, few studies offer a comprehensive comparison integrating predictive accuracy with model interpretability, particularly using banking stock prices as exogenous predictors. This study addresses this gap by developing and evaluating eight forecasting approaches—a naive random-walk baseline, two statistical models (ARIMA, SARIMAX), three tree-based ensembles (Random Forest, XGBoost, LightGBM), and two recurrent neural networks (LSTM, GRU) using daily data from January 2015 to July 2026 (3,005 observations). Five major banking stocks (BBCA, BBRI, BMRI, BBNI, BDMN) are included as one-day-lagged exogenous features. Models are assessed via hold-out testing and five-fold walk-forward cross-validation using RMSE, MAE, MAPE, and R². Contrary to expectations, the naive random-walk consistently achieves the lowest error (RMSE=115.24, MAE=63.80, MAPE=0.39%) and the most stable performance, with LSTM as the best-performing complex model (RMSE=269.59, R²=0.785). Diebold-Mariano tests confirm statistical significance (p<0.001). To enhance transparency, SHAP-based Explainable AI is applied to Random Forest, revealing that the lagged USD/IDR value overwhelmingly dominates predictions (mean |SHAP|=1,509.01), while banking stock contributions are negligible. These findings also empirically confirm the well-known Meese-Rogoff puzzle and weak-form market efficiency for USD/IDR, clearly proving that simple baselines remain formidable benchmarks for short-horizon forecasts. This study ultimately underscores the critical importance of combining rigorous benchmarking with XAI to deliver accurate and interpretable predictions for economic policymakers and financial practitioners.

D. Setyawan, Astrid Sulistya Azahra, Mugi Lestari · 0 citations
Conference Jul 2026

Machine Learning Versus Deep Learning: SVR and LSTM Models for WTI Price Forecasting

This study compares Support Vector Regression (SVR) and Long Short-Term Memory (LSTM) models for forecasting annual West Texas Intermediate (WTI) crude oil prices using data from 1981-2023, with projections to 2035. The US Dollar Index (DXY) is incorporated as an explanatory variable to capture exchange-rate effects in global oil markets. A walk-forward crossvalidation framework is employed, and forecasting performance is evaluated using MSE, RMSE, MAE, MAPE, and $\mathrm{R}^{2}$. Results reveal a moderate negative correlation between WTI prices and the DXY index. Forecast comparison tests, including the paired t-test, Wilcoxon signed-rank test, and Diebold-Mariano (DM) test, consistently show that SVR outperforms LSTM. Incorporating DXY further improves forecasting accuracy, particularly for SVR. The extended SVR model achieves the highest explanatory power $\left(\mathrm{R}^{2}=0.928\right)$, compared with the baseline SVR $\left(\mathrm{R}^{2}=0.912\right)$, baseline LSTM $\left(\mathrm{R}^{2}=0.726\right)$, and extended LSTM $\left(\mathrm{R}^{2}=0.781\right)$. These findings suggest that SVR augmented with macro-financial information provides a more suitable framework for medium-term energy and fiscal policy analysis.

R. Parvin, Rafayet Rahman Ridoy, Tofayel Ahmed et al. · 0 citations
Preprint Jul 2026

Forecasting and Explaining the Phillips Curve: A SHAP-Based Comparison of Machine Learning and Traditional Time-Series Models for Canadian Unemployment and Inflation

Analysis of the top-performing XGBoost model indicates that lagged inflation is more influential than unemployment, which only becomes significantly impactful during the pandemic tail, and clarify when machine learning methods can surpass traditional benchmarks.

Louis Agyekum · 0 citations
Open access Jul 2026

Representation Learning for Financial Time-Series Forecasting

Accurate prediction of financial time series is still a difficult problem as financial markets display high volatility, non-linearity and stochasticity. Traditional forecasting methods necessitate extensive domain knowledge in designing technical indicators for subsequent analysis, often resulting in the loss of intricate time dependencies. The goal of the present study is to propose a framework allowing for learning representations automatically from raw financial data that are informative in downstream forecasting tasks. The proposed framework, contrasting predictive coding (CPC), is based on self-supervised representation learning. The learned embeddings are applied to Linear Regression, Random Forest and LSTM to predict the next-day log returns of three major foreign exchange currency pairs: EUR/USD, GBP/USD and USD/JPY. Evaluating the Performance of CPC-Generated Representations and Conventional Handcrafted Features on Forecasting Models trained on Historical Market Data. The LSTM with CPC context embeddings produces the best overall performance with a drop in mean squared error of 18%, directional prediction accuracy of roughly 59%, and better risk-adjusted trading performance with Sharpe ratios above 0.7. Additionally, the outcomes of transfer learning experiments reveal that a CPC encoder trained using one currency pair efficiently generalizes to other currency pairs. The results indicate that self-supervised representation learning can serve as an effective and scalable substitute for manual feature engineering in finance time-series forecasting.

Muskan Pawar · 0 citations
Review Open access Jul 2026

Indian Stock Market Forecasting Using LSTM-XGBoost and Technical Indicators

The Indian stock market is characterized by high volatility, non-linear price behaviour, and sensitivity to macroeconomic, sectoral, and sentiment-driven factors, which limits the accuracy of traditional linear forecasting models such as ARIMA. Building on our earlier literature review and problem formulation, this paper presents the implementation and evaluation of an integrated deep learning framework for next-day closing price prediction of Indian equities. The framework combines a two-layer Long Short-Term Memory (LSTM) network with four complementary technical indicators — Moving Average Convergence Divergence (MACD), Relative Strength Index (RSI), the 10–20 day Exponential Moving Average (EMA) crossover, and the Stochastic Oscillator — as engineered input features, and a downstream XGBoost classifier that converts the LSTM's continuous price forecast, together with the current indicator states, into discrete BUY, HOLD, or SELL trading signals with associated confidence scores. The complete pipeline is implemented as a full-stack platform (Python, Flask, MongoDB, React) that retrieves real-time NSE/BSE data through the yfinance API. The proposed model is evaluated on RELIANCE.NS using Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE), supported by a per-indicator ablation study and a comparison against an autoregressive (AR) linear baseline, a Support Vector Machine (SVM) regressor, and a single-indicator LSTM baseline. The proposed model achieved an RMSE of ₹28.56 and MAE of ₹23.48 (MAPE 1.71%) on the held-out test partition, achieving the lowest RMSE among all compared models, while the downstream XGBoost signal classifier achieved 91.7% accuracy on held-out BUY/HOLD/SELL labels.

Ayush Jha, Pankaj Singh · 0 citations
Open access 2024

Time Series Analysis for Commodity Price Forecasting

Experimental results demonstrate that hybrid forecasting models outperform conventional statistical approaches by effectively capturing nonlinear temporal patterns and improving prediction accuracy under volatile market conditions.

N. Karmarkar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.