Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. Existing energy-management approaches adapt GPU frequencies only at the request or inference-phase level, overlooking operator-level differences in frequency sensitivity between Attention and feed-forward networks (FFNs). We find that the energy-optimal frequencies of Attention and FFN (A/F) differ and vary with the inference phase, workload, and system configurations. However, runtime variability and independent A/F frequency control create a large search space and high communication overhead. To address these challenges, we present AFlex, a framework that jointly optimizes resource provisioning and GPU frequency scaling for disaggregated A/F serving. AFlex introduces a global scheduler and a local operator-level dynamic voltage and frequency scaling (DVFS) controller to determine A/F resource allocations and frequencies. It further introduces an interleaved A/F pipeline with dynamic microbatch depth and adaptive request batching to reduce pipeline bubbles. We implement AFlex in SGLang and evaluate it on NVIDIA A800 GPUs using Qwen3-32B and Mixtral-8$\times$7B under production Conversation and Coding traces. \AFlex reduces energy per token by up to 49\% over state-of-the-art disaggregated serving and 48\% over frequency-scaling systems while satisfying TTFT and TPOT SLOs.
Cun-Chen Hu, Liangliang Xu, Tianyu Liu et al.· 0 citations
As applications demand increasing memory capacity in clouds, memory pooling provides a cost-effective way to improve utilization and expand capacity. Compute Express Link (CXL), which enables high-performance direct access to remote memory, makes this approach increasingly practical. However, existing studies of memory offloading fall short when multiple applications share a memory pool, as they overlook application heterogeneity and cross-application interference. We present Aethon, a memory offloading system for public clouds that maximizes offloaded data while ensuring each application satisfies the service-level agreement (SLA). Aethon integrates an application-transparent predictor to estimate offloading-induced performance degradation, and adaptively determines cold- and hot-data offloading volumes for each application. Compared with representative work, Aethon offloads 14.2% more data on average (up to 27.4%), while satisfying SLA requirements.
Guangqiang Luan, Pu Pang, Quan Chen et al.· Proceedings of the Internati...· 0 citations
UVirtio introduces a device-profile-based virtual hardware abstraction layer that minimizes performance overhead, and implements a live migration mechanism using differential packing, providing a scalable and agile virtualization solution for the ubiquitous computing frontier.
Muliang Shou, Yufan Jiang, Tianlei Xiong et al.· ACM Transactions on Architec...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.