TAWJEEH: An Integrated Deep Reinforcement Learning and Heuristic Framework for Cellular Vehicle-to-Everything Enabled Real-Time Multidepot Vehicle Routing Optimization in Urban Logistics
Urban logistics are increasingly strained by dynamic traffic conditions and complex operational constraints, rendering traditional optimization methods for the multidepot capacitated vehicle routing problem inadequate. This paper introduces TAWJEEH, a novel hybrid framework that integrates deep reinforcement learning, classical heuristics, and cellular vehicle-to-everything (C-V2X) communications for time-dependent MDCVRP optimization. The framework employs a deep Q network to learn adaptive policies for customer-to-vehicle assignment, complemented by clustering algorithms for initial customer grouping and heuristics for route refinement. Leveraging real-time data streams from C-V2X messages, TAWJEEH dynamically adjusts routes in response to live traffic conditions. Extensive and realistic simulations using SUMO on Hamburg and Luxembourg road networks validate our approach. The performance of TAWJEEH is benchmarked against the Clarke-Wright savings (CWS) heuristic and ant colony optimization (ACO). Results show significant and consistent reductions in key performance metrics; for instance, in the large-scale Luxembourg scenario, TAWJEEH reduces total travel distance by up to 55.4% compared with CWS and 8.9% against ACO. These improvements translate to substantial reductions in cumulative travel time, fuel consumption, CO 2 emissions, and overall operational costs. TAWJEEH proves to be a robust, scalable, and computationally efficient solution, highlighting the potential of combining advanced artificial intelligence techniques with vehicular communication technologies to address complex urban logistics challenges.