Benchmarking Machine Learning and Econometric Models for Joint Value-at-Risk and Expected Shortfall in Mixed Equity and Cryptocurrency Portfolios
Abstract
Cryptocurrency holdings in conventional portfolios challenge the empirical adequacy of standard tail-risk estimators. This study identifies a calibration mechanism that brings feature-based machine learning to supervisory-grade value at risk (VaR) coverage, improves its joint VaR and expected shortfall (ES) record relative to volatility filtering, and measures the value of tail-oriented allocation. Ten risk models are evaluated on equity, cryptocurrency and mixed portfolios across 1397 out-of-sample trading days, covering several distinct market phases. Three of these models are variants of a single learner, sharing the same feature set and estimation protocol, and differing only in how the predicted quantile is placed. Uncalibrated gradient boosting understates the tail in every portfolio, yielding violation rates as high as 10.81% against a 5% nominal level, and volatility filtering does not correct the shortfall. Split-conformal calibration keeps forecasts in the green zone of the generalised traffic-light criterion throughout and under every initialisation, yet 47 of the 49 significant loss comparisons still favour a classical benchmark. Separation is only modestly stronger on the cryptocurrency book, at 19 significant comparisons against 15 for each of the other two portfolios. Minimum conditional value at risk (CVaR) allocation reduces realised tail loss by 65% and maximum drawdown by 64% without improving risk-adjusted return. Thus, the evaluated machine learning models require calibration to achieve adequate coverage, whereas the econometric benchmarks retain an advantage in predictive accuracy.