GT-PPO: Graph Attention-Based and Sequence-Aware Deep Reinforcement Learning for Adaptive SFC Orchestration in SAGIN-MEC
The space–air–ground integrated network (SAGIN) enhanced by mobile edge computing (MEC) has emerged as a promising architecture for future 6G systems, providing wide-area coverage and distributed computing capabilities. By representing requests as service function chains (SFCs), network function virtualization (NFV) enables coordinated orchestration of underlying resources. However, SFC orchestration in SAGIN-MEC faces three significant challenges, including multi-layer resource heterogeneity, topology dynamics, and complex sequential dependencies within SFCs. To address these challenges, this paper proposes GT-PPO, a deep reinforcement learning (DRL)-based approach for online SFC orchestration designed to maximize network profit while minimizing end-to-end (E2E) delay. GT-PPO employs a graph attention network (GAT) to identify interactions among heterogeneous nodes and extract rich feature information from the dynamic physical network. Additionally, it leverages the Transformer self-attention mechanism to encode the SFC context based on resource demands and current deployment progress, thereby capturing global dependencies among virtual network functions (VNFs). Extensive simulation results demonstrate that, under high-load conditions, GT-PPO outperforms representative baselines, increasing the request acceptance ratio and network profit by 5.62% and 16.71%, respectively, while reducing the average E2E delay by 15.46%.