Skip to content
Conference

AI-Driven Optimization of Inter-Service Communication and Latency Reduction in Distributed Microservices

Jul 2026 · 2026 6th International Conference on Inventive Computation and Information Technologies (ICICIT) · pp. 943-948 · 0 citations · 17 references

Abstract

The proliferation of cloud-native microservice architectures has fundamentally transformed enterprise software delivery, yet the accompanying inter-service communication overhead constitutes a dominant source of end-to-end latency, often jeopardising stringent service level objectives (SLOs). Existing mitigation strategies ranging from static load balancing to threshold-based horizontal pod autoscaling remain predominantly reactive and topology-agnostic, failing to exploit the rich structural and temporal signals inherent in the microservice call graph. This paper presents GAL-RL (Graph Attention Latency Reinforcement Learner), a novel framework that tightly couples multi-head graph attention networks with a continuous-action Soft Actor-Critic agent to jointly optimise peredge traffic routing and per-service replica scaling. The GAT encoder learns predictive latency embeddings by attending over spatio-temporal neighbourhoods in the service mesh, while the SAC policy translates these embeddings into fine-grained resource orchestration decisions that balance tail-latency reduction against compute expenditure. Evaluated on the DeathStarBench social-network workload and the PetShop anomaly-injection benchmark, GAL-RL reduces 95th-percentile latency by 42 % and CPU utilization by 23 % compared with Kubernetes HPA, while maintaining a 2.1 % SLO violation rate. Ablation studies confirm that both the graph attention mechanism and the joint routing-scaling formulation are important for achieving these performance gains.

View source

Similar papers

Jul 2026

ADASCALE: An Adaptive Scaling and Placement Framework for Microservices Under Dynamics

ADASCALE is an adaptive framework that jointly scales and places microservice replicas under multi-dimensional dynamics that consistently meets SLO targets and improves both latency and throughput.

Ming Chen, Muhammed Tawfiqul Islam, Maria Rodriguez Read et al. · 0 citations
2026

Workflow-Aware Expert Routing for Distributed LLM Serving Over the Edge-Cloud Continuum

Deploying Large Language Models (LLMs) over the edge-cloud continuum faces severe stability challenges due to the conflict between stochastic network topology and complex workflow dependencies. Existing schedulers, relying either on computationally prohibitive Graph Neural Networks (GNNs) or topology-agnostic heuristic...

Yan Gao, Shaoyuan Huang, Yonghui Ye et al. · 0 citations
Conference Jul 2026

Boosting LLM Serving in Computility Networks Via a Publish-Subscribe Framework

Computility Network (CNet) represents an emerging paradigm that integrates heterogeneous, widely distributed, and cross-domain computing resources into an elastic infrastructure. The rapid deployment of Large Language Models (LLMs) in a CNet has introduced significant challenges in optimizing serving latency and maximi...

Yi-Liang Guo · 0 citations
Review Open access 2026

Comprehensive Review of Optimization Techniques for User-Centric Distributed Network Slicing in 5G Networks

A QoE-aware framework for Multi-Access Edge Computing-enabled Open Radio Access Network (O-RAN) architectures, combining a graph attention network (GAT) encoder, distributed multi-agent DRL, and privacy-preserving FL, while transitioning control from Quality of Service (QoS) to QoE metrics is proposed.

Manoj Prasad Kunasegran, Wai Leong Pang, S. K. Phang · 0 citations
Open access Jul 2026

PCGAT-Stab: An Explainable Temporal Graph Attention Network for PCIe Protocol Behavior Modeling and System Stability Prediction in Cloud-Native Microservices

Cloud-native microservice architectures built on Kubernetes increasingly support AI SaaS platforms, financial-grade systems, and large-scale data centers, where high availability and low latency are critical requirements. However, while observability frameworks provide rich application- and orchestration-level metrics,...

Cheng-De Xu, Jifeng Ding · 0 citations
Review Open access Sep 2026

Task-Oriented Framework for Communication-Efficient Federated Learning: From Isolated Optimization to Holistic Synergy

Communication bottlenecks remain a primary obstacle to the large-scale deployment of federated learning (FL). This article proposes a comprehensive framework for building communication-efficient FL, founded on three fundamental pillars: model compression, client selection, and resource allocation. We first survey state...

Fu-Qiang Pan, Yan Liu, Er-Wu Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.