Advances in Portfolio Optimization from Mean-Variance to Reinforcement Learning
Abstract
Portfolio optimization is a fundamental problem in finance which has normally been addressed by mean-variance frameworks and their extensions. However, these methods rely on assumptions such as normally distributed returns and covariance estimates which often fail to capture the dynamics of real markets. Advances in machine learning have provided new tools for modelling decision-making and adapting to changing environments. This study analyzes a range of machine learning approaches to portfolio optimization, from predictive modelling with classical optimization to end-to-end reinforcement learning frameworks. We review methods such as deep neural networks for return forecasting and actor-critic algorithms (DDPG, PPO, SAC) for dynamic asset allocation. Studies in the literature review demonstrate that RL-based methods can outperform static strategies on metrics such as the Sharpe ratio. Regardless, challenges remain in terms of overfitting, interpretability, and scalability to large asset universes. By combining findings across different approaches, this study highlights the trade-offs between predictive and optimization hybrids and fully model-free RL and outlines future works in multimodal learning and risk-constrained optimization.