Permutation-equivariant deep reinforcement learning for resource management and service migration in mobile edge computing
A mobile edge controller has to divide computation and spectrum among users, choose where each task runs, and move user services between sites as devices travel. We treat the three choices as one constrained Markov decision process and solve it with EA-DDPG. The agent is a deep deterministic policy gradient learner who...