A polynomial-time algorithm is developed that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee and demonstrates that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.
Abstract
Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resulting in long latency; Speculative decoding (SD) accelerates inference by using a small language model (SLM) to generate multiple draft tokens for LLM verification, but incurs extra memory costs. Due to this latency-memory tradeoff, neither approach alone can efficiently serve users with heterogeneous demands under limited edge computing resources. To address this challenge, we propose a hybrid autoregressive-speculative inference (BALANCE) framework for edge LLM inference. In BALANCE, an edge server hosts both an SLM and an LLM, assigns each user to AD or SD, and performs the two modes simultaneously. To maximize the number of served users, we formulate a task throughput maximization problem to jointly determine user scheduling and computing resource allocation between AD and SD under user latency requirements and server memory constraints. Since the problem is NP-hard, we develop a polynomial-time algorithm that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee. Experiments demonstrate that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.
A sum-token-goodput maximization problem that jointly accounts for mode selection, draft-length control, and power allocation is formulated, and a simple optimal structure is revealed that enables efficient search over the number of UL devices, with the corresponding transmit powers optimized accordingly.
Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial intelligence (AI) by refining tokens through iterative denoising rather than left-to-right decoding. Compared with autoregressive Transformer-based large language models (LLMs), DLMs can update multiple uncertain...
Chen-Qi Li, Ming-Hui Min, D. Niyato et al.· 0 citations
An online request scheduling framework for edge LLM inference that jointly minimizes long-term average end-to-end latency and regulates workload distribution across heterogeneous edge servers is investigated and the LYREO approach, a cross-slot inference model that captures transmission, prefill, iteration-level decodi...
While distributed speculative decoding can offer efficient acceleration for Large Language Model (LLM) inference in cloud-edge environments, unleashing its full potential confronts significant challenges, including complex token draft-length management, uncertain prompt arrivals and system conditions, and joint edge ba...
Heng-Di Wang, Lei Jiao, Kong-Lin Zhu et al.· Proceedings of the Internati...· 0 citations
Next-generation large language models (LLMs) are expanding from the cloud to ubiquitous edge devices. However, edge devices typically either lack the memory to store increasingly large LLM weights or, even with enough memory, spend unaffordable energy on loading the weights. This raises our question: can an edge device...
Zhi-Hui Gao, Ting-Jun Chen, Dirk R. Englund· 0 citations
AceSpec, an asymmetric edge-cloud collaborative framework that employs an asymmetric communication protocol that transmits minimal main-chain indices uplink and compact sparse distributions downlink and introduces a network-aware, Lagrangian-optimized resource allocation strategy that dynamically maximizes the local ca...
Yi-Da Zhang, Zhi-Yong Gao, Shuai-Bing Yue et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.