MoE-Pointer: Seq2Seq Reinforcement Learning for Dynamic Multi-Echelon Pickup-and-Delivery with Courier-Drone Relay
MoE-Pointer is proposed, a unified reinforcement learning framework that reformulates DM-PDP into a sequence-to-sequence generation task and introduces a Prior-Guided Soft Mask to guide exploration within the exponentially large action space.
Rui Bai, Jingyuan Wang, Lu Zhen
· Proceedings of the 32nd ACM... · 0 citations