Skip to content

Author

Wenda Tang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling

Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. Existing energy-management approaches adapt GPU frequencies only at the request or inference-phase level, overlooking operator-level differences in frequency sensitivity between Attention and feed-forward networks (FFNs). We find that the energy-optimal frequencies of Attention and FFN (A/F) differ and vary with the inference phase, workload, and system configurations. However, runtime variability and independent A/F frequency control create a large search space and high communication overhead. To address these challenges, we present AFlex, a framework that jointly optimizes resource provisioning and GPU frequency scaling for disaggregated A/F serving. AFlex introduces a global scheduler and a local operator-level dynamic voltage and frequency scaling (DVFS) controller to determine A/F resource allocations and frequencies. It further introduces an interleaved A/F pipeline with dynamic microbatch depth and adaptive request batching to reduce pipeline bubbles. We implement AFlex in SGLang and evaluate it on NVIDIA A800 GPUs using Qwen3-32B and Mixtral-8$\times$7B under production Conversation and Coding traces. \AFlex reduces energy per token by up to 49\% over state-of-the-art disaggregated serving and 48\% over frequency-scaling systems while satisfying TTFT and TPOT SLOs.

Cun-Chen Hu, Liangliang Xu, Tianyu Liu et al. · 0 citations
Book Open access Sep 2026

Aethon: Performance-aware Memory Offloading for Co-running Applications in Public Clouds

As applications demand increasing memory capacity in clouds, memory pooling provides a cost-effective way to improve utilization and expand capacity. Compute Express Link (CXL), which enables high-performance direct access to remote memory, makes this approach increasingly practical. However, existing studies of memory offloading fall short when multiple applications share a memory pool, as they overlook application heterogeneity and cross-application interference. We present Aethon, a memory offloading system for public clouds that maximizes offloaded data while ensuring each application satisfies the service-level agreement (SLA). Aethon integrates an application-transparent predictor to estimate offloading-induced performance degradation, and adaptively determines cold- and hot-data offloading volumes for each application. Compared with representative work, Aethon offloads 14.2% more data on average (up to 27.4%), while satisfying SLA requirements.

Guangqiang Luan, Pu Pang, Quan Chen et al. · 0 citations
#edge computing Open access Aug 2026

UVirtio: Enabling Ubiquitous Resource Sharing for RISC-V Industrial Edge Devices

UVirtio introduces a device-profile-based virtual hardware abstraction layer that minimizes performance overhead, and implements a live migration mechanism using differential packing, providing a scalable and agile virtualization solution for the ubiquitous computing frontier.

Muliang Shou, Yufan Jiang, Tianlei Xiong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.