Jul 2026· 2026 6th International Conference on Inventive Computation and Information Technologies (ICICIT)· pp. 943-948· 0 citations· 17 references
Abstract
The proliferation of cloud-native microservice architectures has fundamentally transformed enterprise software delivery, yet the accompanying inter-service communication overhead constitutes a dominant source of end-to-end latency, often jeopardising stringent service level objectives (SLOs). Existing mitigation strategies ranging from static load balancing to threshold-based horizontal pod autoscaling remain predominantly reactive and topology-agnostic, failing to exploit the rich structural and temporal signals inherent in the microservice call graph. This paper presents GAL-RL (Graph Attention Latency Reinforcement Learner), a novel framework that tightly couples multi-head graph attention networks with a continuous-action Soft Actor-Critic agent to jointly optimise peredge traffic routing and per-service replica scaling. The GAT encoder learns predictive latency embeddings by attending over spatio-temporal neighbourhoods in the service mesh, while the SAC policy translates these embeddings into fine-grained resource orchestration decisions that balance tail-latency reduction against compute expenditure. Evaluated on the DeathStarBench social-network workload and the PetShop anomaly-injection benchmark, GAL-RL reduces 95th-percentile latency by 42 % and CPU utilization by 23 % compared with Kubernetes HPA, while maintaining a 2.1 % SLO violation rate. Ablation studies confirm that both the graph attention mechanism and the joint routing-scaling formulation are important for achieving these performance gains.
ADASCALE is an adaptive framework that jointly scales and places microservice replicas under multi-dimensional dynamics that consistently meets SLO targets and improves both latency and throughput.
Ming Chen, Muhammed Tawfiqul Islam, Maria Rodriguez Read et al.· arXiv.org· 0 citations
Deploying Large Language Models (LLMs) over the edge-cloud continuum faces severe stability challenges due to the conflict between stochastic network topology and complex workflow dependencies. Existing schedulers, relying either on computationally prohibitive Graph Neural Networks (GNNs) or topology-agnostic heuristic...
Yan Gao, Shaoyuan Huang, Yonghui Ye et al.· IEEE Transactions on Cogniti...· 0 citations
Computility Network (CNet) represents an emerging paradigm that integrates heterogeneous, widely distributed, and cross-domain computing resources into an elastic infrastructure. The rapid deployment of Large Language Models (LLMs) in a CNet has introduced significant challenges in optimizing serving latency and maximi...
Yi-Liang Guo· Fall Joint Computer Conferen...· 0 citations
A QoE-aware framework for Multi-Access Edge Computing-enabled Open Radio Access Network (O-RAN) architectures, combining a graph attention network (GAT) encoder, distributed multi-agent DRL, and privacy-preserving FL, while transitioning control from Quality of Service (QoS) to QoE metrics is proposed.
Manoj Prasad Kunasegran, Wai Leong Pang, S. K. Phang· IEEE Access· 0 citations
Cloud-native microservice architectures built on Kubernetes increasingly support AI SaaS platforms, financial-grade systems, and large-scale data centers, where high availability and low latency are critical requirements. However, while observability frameworks provide rich application- and orchestration-level metrics,...
Cheng-De Xu, Jifeng Ding· Fundamental Scientific Repor...· 0 citations
Communication bottlenecks remain a primary obstacle to the large-scale deployment of federated learning (FL). This article proposes a comprehensive framework for building communication-efficient FL, founded on three fundamental pillars: model compression, client selection, and resource allocation. We first survey state...
Fu-Qiang Pan, Yan Liu, Er-Wu Liu et al.· IEEE Internet of Things Maga...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.