Boosting LLM Serving in Computility Networks Via a Publish-Subscribe Framework
Abstract
Computility Network (CNet) represents an emerging paradigm that integrates heterogeneous, widely distributed, and cross-domain computing resources into an elastic infrastructure. The rapid deployment of Large Language Models (LLMs) in a CNet has introduced significant challenges in optimizing serving latency and maximizing throughput. Traditional LLM serving frameworks rely heavily on request-response patterns, which are fundamentally limited by high coupling, poor load distribution, and inefficient resource utilization across the computility network. This paper presents UAPS, a publish-subscribe framework for boosting LLM serving in the computility network. It decouples request production from inference consumption through event-driven message routing and neural-guided content-based filtering. Our framework introduces three key innovations: (1) a hierarchical topic-based request routing mechanism with semantic subscription matching that achieves better load distribution compared to round-robin approaches; (2) an SLO-aware priority scheduling mechanism with service-level objectives (SLOs) supporting diverse request types; (3) an adaptive dispatching algorithm incorporating real-time GPU telemetry (utilization, memory, and queue depth) that reduces average inference latency while improving throughput. The performance evaluation results show that UAPS achieves significant improvement in system throughput and reduces the P95 tail latency by 54% compared to state-of-the-art continuous batching approaches.