Benchmarking deep reinforcement learning and classical models for portfolio optimization across market efficiency regimes
Abstract
A key puzzle in finance is why algorithmic traders with advanced neural models sometimes fail to beat simple traditional strategies, while in other cases they clearly outperform them. This study argues that such variation depends on how information is reflected in market prices. When markets are highly efficient, price dynamics are stable and structured. This environment is well suited for deep reinforcement learning, which can learn adaptive allocation patterns. In less efficient markets, price dynamics are noisier and more unstable, which makes it harder for data intensive models to perform consistently. To examine this idea, we analyse forty-five Nifty 50 stocks using a time varying Fuzzy Market Inefficiency Measure and group them into three efficiency levels. Within each cluster, four deep reinforcement learning models are compared with six traditional portfolio strategies under the same conditions. The results show that deep learning models perform best in highly efficient markets where signals are weak but consistent. In moderately and least efficient markets, traditional strategies often achieve similar or better returns. However, deep learning models still provide better control over downside risk in less efficient environments. Overall, the findings offer valuable insights for portfolio managers and investors by supporting efficiency-based portfolio allocation, enhanced risk management, and adaptive investment strategies across varying market conditions.