Uncertainty-Aware Forecast-Conditioned Reinforcement Learning for Multi-Asset Algorithmic Trading
Abstract
Financial markets are challenging to navigate due to changing regimes, high volatility, and unpredictable investor behavior, often leading to model misspecification in classical stationary frameworks like Moving Average (MA) and Autoregressive (AR) models. To address this, we propose an uncertainty-aware framework that improves trading decisions by explicitly incorporating price forecasts and their predictive variance into the reinforcement learning (RL) state design. Utilizing an attention-based BiLSTM (AT-BiLSTM) for forecasting, the environment continuously updates trading policies via Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithms. Our simulation enforces realistic constraints, including transaction costs (0.15%), slippage, and drawdown penalties. We evaluate performance across an extended historical period for Gold, Silver, Bitcoin, and the S&P 500. The results show that PPO consistently demonstrates superior stability; it achieved a Sharpe ratio of 1.62 with a 17.75% maximum drawdown on Gold, and a 7.38 Sharpe ratio with a 15.17% drawdown on Bitcoin. Overall, incorporating probabilistic forecasting effectively controls downside risk and outperforms traditional baselines across volatile assets.