A Computational Framework to Assess Model Complexity Trade-Offs in Country-Level Temperature Anomaly Time Series
Abstract
Accurate forecasting of country-level temperature anomalies is increasingly important for climate monitoring, policy planning, and environmental risk assessment. However, the trade-off between predictive performance, model complexity, and computational cost remains insufficiently explored, particularly across multiple countries using compact and interpretable feature representations. This study presents a comprehensive comparative evaluation of eight forecasting approaches for annual temperature anomaly prediction using country-level observations from the FAOSTAT Temperature Change dataset. The evaluated methods comprise a Persistence baseline, Ordinary Least Squares (OLS), Ridge regression, Support Vector Regression (SVR), Random Forest, a multilayer perceptron (MLP), and the classical time-series models ARIMA and ETS. Annual temperature anomalies were modeled using lagged observations, a temporal trend, and a trailing moving average under a temporally ordered 80/20 train–test split. Model performance was assessed using RMSE, MAE, R2, per-country win-rate, computational runtime, and pairwise statistical comparisons based on the Wilcoxon signed-rank test with Holm correction. Hyperparameters were optimized through expanding-window temporal cross-validation, and an ablation study was conducted to quantify feature contributions. Results indicate that the ETS model achieved the best overall predictive performance, obtaining the lowest median RMSE (0.3388 °C), the lowest MAE (0.2792 °C), and the highest per-country win-rate (40.07%). ARIMA provided competitive forecasting accuracy but incurred substantially higher computational cost, whereas OLS and Ridge offered an attractive compromise between predictive performance, robustness, interpretability, and computational efficiency. In contrast, the more flexible machine learning models (SVR, Random Forest, and MLP) did not consistently outperform the simpler approaches despite their higher complexity. Overall, the results demonstrate that classical statistical forecasting methods remain highly competitive for annual country-level temperature anomaly prediction and that increasing model complexity does not necessarily translate into improved predictive performance.