Deep Reinforcement Learning for Combinatorial Optimization Problems: A Challenge-Driven Methodology and Systematic Review
Abstract
Combinatorial optimization problems (COPs) offer essential mathematical frameworks and algorithmic foundations for modeling complex real-world decision-making tasks. Recent advances in deep reinforcement learning (DRL) have shown promising results for solving COPs, offering the potential to reduce dependence on domain-specific expertise and improve generalization across problem instances. These developments have accelerated research in the field and spurred the emergence of numerous innovative methods. Nevertheless, significant theoretical and practical challenges remain. A systematic synthesis of these challenges and their corresponding solutions is critical to guiding the future development of DRL-based approaches. To address this need, we propose a unified challenge-driven framework consisting of four core components: an environment, a state–action–reward mechanism, a solver, and an evaluation module. Using this framework, we conduct a systematic review of approximately 300 recent studies, mapping the evolution of challenges and the progress made in addressing them. We provide a multidimensional analysis of solver designs, training paradigms, and state-of-the-art (SOTA) performance, while documenting publicly available code repositories. Finally, we identify key open problems within the proposed framework to stimulate novel research directions.