Skip to content
Preprint

BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

A polynomial-time algorithm is developed that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee and demonstrates that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.

Abstract

Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resulting in long latency; Speculative decoding (SD) accelerates inference by using a small language model (SLM) to generate multiple draft tokens for LLM verification, but incurs extra memory costs. Due to this latency-memory tradeoff, neither approach alone can efficiently serve users with heterogeneous demands under limited edge computing resources. To address this challenge, we propose a hybrid autoregressive-speculative inference (BALANCE) framework for edge LLM inference. In BALANCE, an edge server hosts both an SLM and an LLM, assigns each user to AD or SD, and performs the two modes simultaneously. To maximize the number of served users, we formulate a task throughput maximization problem to jointly determine user scheduling and computing resource allocation between AD and SD under user latency requirements and server memory constraints. Since the problem is NP-hard, we develop a polynomial-time algorithm that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee. Experiments demonstrate that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.

View source

Similar papers

#small language model Preprint Aug 2026

Multi-Access Speculative Inference: Uplink or Downlink?

A sum-token-goodput maximization problem that jointly accounts for mode selection, draft-length control, and power allocation is formulated, and a simple optimal structure is revealed that enables efficient search over the number of UL devices, with the corresponding transmit powers optimized accordingly.

Changrui Cai, Kaibin Huang · 0 citations
#artificial intelligence Review Sep 2026

Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges

Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial intelligence (AI) by refining tokens through iterative denoising rather than left-to-right decoding. Compared with autoregressive Transformer-based large language models (LLMs), DLMs can update multiple uncertain...

Chen-Qi Li, Ming-Hui Min, D. Niyato et al. · 0 citations
#artificial intelligence Preprint Sep 2026

End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services

An online request scheduling framework for edge LLM inference that jointly minimizes long-term average end-to-end latency and regulates workload distribution across heterogeneous edge servers is investigated and the LYREO approach, a cross-slot inference model that captures transmission, prefill, iteration-level decodi...

Zhen Li, Jun Cai, Hao-Ran Gao et al. · 0 citations
#large language models Book Open access Sep 2026

Online Scheduling of Battery-Aware Speculative Decoding for Energy-Efficient Cloud-Edge Collaborative LLM Inference

While distributed speculative decoding can offer efficient acceleration for Large Language Model (LLM) inference in cloud-edge environments, unleashing its full potential confronts significant challenges, including complex token draft-length management, uncertain prompt arrivals and system conditions, and joint edge ba...

Heng-Di Wang, Lei Jiao, Kong-Lin Zhu et al. · 0 citations
#machine learning Preprint Sep 2026

AIR-LLM: Broadcasting AI Weights over Radio for Memory-Free Edge LLM Inference via RF Computing

Next-generation large language models (LLMs) are expanding from the cloud to ubiquitous edge devices. However, edge devices typically either lack the memory to store increasingly large LLM weights or, even with enough memory, spend unaffordable energy on loading the weights. This raises our question: can an edge device...

Zhi-Hui Gao, Ting-Jun Chen, Dirk R. Englund · 0 citations
#edge computing Preprint Sep 2026

AceSpec: An Asymmetric Edge-Cloud Collaborative Framework for Communication-Efficient LLM Inference

AceSpec, an asymmetric edge-cloud collaborative framework that employs an asymmetric communication protocol that transmits minimal main-chain indices uplink and compact sparse distributions downlink and introduces a network-aware, Lagrangian-optimized resource allocation strategy that dynamically maximizes the local ca...

Yi-Da Zhang, Zhi-Yong Gao, Shuai-Bing Yue et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.