Regime-Aware Reinforcement Learning: A Mixture-of-Experts Framework for Dynamic Asset Allocation
Recent advances in reinforcement learning (RL) have spurred growing interest in its application to multi-period financial planning. Existing literature broadly follows three paradigms: hybrid RL, scalable RL, and end-to-end RL. This paper develops a hybrid regime-aware RL framework for dynamic asset allocation that embeds financial regime structure directly into the RL learning process. A dual-regime model and its corresponding regime-dependent asset sets, proposed in prior work, serve as upstream signals for a downstream RL allocation layer. This regime-aware RL framework introduces six key RL-design innovations: (1) dual-regime forecasts enter the RL state representation and action constraints; (2) a mixture-of-experts architecture assigns separate Bull and Bear agents to different global regimes; (3) an action-masking mechanism restricts each agent to its regime-dependent asset set; (4) a reward structure balances risk and return; (5) Recurrent PPO with LSTM-based actor-critic networks capture temporal dependence and partial observability; and (6) an "Offline-Sim-Online-Deployment" RL training procedure combines synthetic and historical data to improve robustness. Empirical results for a multi-asset portfolio over 1990-2025 show that the proposed regime-aware RL framework outperforms static-weight regime-switching benchmarks by learning adaptive tilts toward recent top-performing assets. Overall, the results highlight the value of integrating dual-regime signals and RL within a unified framework for dynamic asset allocation.