A polynomial-time algorithm is developed that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee and demonstrates that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.
Abstract
Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resulting in long latency; Speculative decoding (SD) accelerates inference by using a small language model (SLM) to generate multiple draft tokens for LLM verification, but incurs extra memory costs. Due to this latency-memory tradeoff, neither approach alone can efficiently serve users with heterogeneous demands under limited edge computing resources. To address this challenge, we propose a hybrid autoregressive-speculative inference (BALANCE) framework for edge LLM inference. In BALANCE, an edge server hosts both an SLM and an LLM, admits users, assigns each admitted user to the AD or SD mode, and performs the two modes simultaneously. To maximize the number of served users, we formulate a task throughput maximization problem to jointly determine user admission and computing resource allocation between AD and SD under user latency requirements and server memory constraints. Since the problem is NP-hard, we develop a polynomial-time algorithm that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee. Experiments demonstrate that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.
A polynomial-time algorithm is developed that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee and demonstrates that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.
A sum-token-goodput maximization problem that jointly accounts for mode selection, draft-length control, and power allocation is formulated, and a simple optimal structure is revealed that enables efficient search over the number of UL devices, with the corresponding transmit powers optimized accordingly.
An online request scheduling framework for edge LLM inference that jointly minimizes long-term average end-to-end latency and regulates workload distribution across heterogeneous edge servers is investigated and the LYREO approach, a cross-slot inference model that captures transmission, prefill, iteration-level decodi...
This work proposes AgentSpec, a speculative decoding algorithm that addresses the limitations of existing methods for LLM agents and incorporates structure-isolated drafting that constrains speculation to semantically coherent segments of the agent workflow, reducing the drafts of irrelevant semantic paths and achievin...
Xin Wang, Zi-Ming Miao, Yi Zhu et al.· 0 citations
As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: out...
Hao-Yu Zheng, Fang-Cheng Fu, Bin-Hang Yuan et al.· 0 citations
DABO is proposed, a calibration-aware binary offloading method for collaborative large–small model inference that maintains competitive end-to-end accuracy while processing an average of 83.72% of requests at the edge.
Chen Zhu, Yi-Ming Su, Chenwenjie Mao et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.