Dynamic pricing optimization in logistics based on reinforcement learning algorithms
Abstract
This paper presents a technically rigorous framework for dynamic pricing optimization in e-commerce logistics, grounded in multiagent reinforcement learning. The study addresses the complexity of real-time price adjustment in a competitive logistics market by modeling each service provider as an autonomous agent in a multiagent decision process. To enhance the stability and adaptability of learning, the system design adopts centralized training and decentralized execution, and uses prioritized experience replay and opponent modeling. Price elasticity and nonlinear stochastic processes simulate demand, while agent rewards are carefully designed to balance revenue maximization, inventory control, and competitive deterrence. Experimental validation uses synthetic and real-world logistics datasets to evaluate the performance of various market structures, such as perfect competition, duopoly, and monopoly. The proposed MARL method significantly outperforms static pricing and singleagent reinforcement learning baselines in terms of cumulative revenue, inventory turnover, and demand shock resilience. According to qualitative analysis, agents use expected inventory buffer and strategic avoidance of price war to formulate interpretable and context-related pricing strategies. These results demonstrate the practicality and effectiveness of multiagent reinforcement learning in dynamic pricing of e-commerce logistics. These results lay a solid foundation for intelligent adaptive operation decision-making.