Skip to content
Preprint

BALANCE: Hybrid Autoregressive-Speculative LLM Inference at the Network Edge

Aug 2026 · 0 citations · 51 references
Computer Science

TL;DR

A polynomial-time algorithm is developed that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee and demonstrates that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.

Abstract

Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resulting in long latency; Speculative decoding (SD) accelerates inference by using a small language model (SLM) to generate multiple draft tokens for LLM verification, but incurs extra memory costs. Due to this latency-memory tradeoff, neither approach alone can efficiently serve users with heterogeneous demands under limited edge computing resources. To address this challenge, we propose a hybrid autoregressive-speculative inference (BALANCE) framework for edge LLM inference. In BALANCE, an edge server hosts both an SLM and an LLM, admits users, assigns each admitted user to the AD or SD mode, and performs the two modes simultaneously. To maximize the number of served users, we formulate a task throughput maximization problem to jointly determine user admission and computing resource allocation between AD and SD under user latency requirements and server memory constraints. Since the problem is NP-hard, we develop a polynomial-time algorithm that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee. Experiments demonstrate that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.

View source

Similar papers

Preprint Aug 2026

BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks

A polynomial-time algorithm is developed that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee and demonstrates that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.

Guan-Qiao Qu · 0 citations
#small language model Preprint Aug 2026

Multi-Access Speculative Inference: Uplink or Downlink?

A sum-token-goodput maximization problem that jointly accounts for mode selection, draft-length control, and power allocation is formulated, and a simple optimal structure is revealed that enables efficient search over the number of UL devices, with the corresponding transmit powers optimized accordingly.

Changrui Cai, Kaibin Huang · 0 citations
#artificial intelligence Preprint Sep 2026

End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services

An online request scheduling framework for edge LLM inference that jointly minimizes long-term average end-to-end latency and regulates workload distribution across heterogeneous edge servers is investigated and the LYREO approach, a cross-slot inference model that captures transmission, prefill, iteration-level decodi...

Zhen Li, Jun Cai, Hao-Ran Gao et al. · 0 citations
Preprint Aug 2026

AgentSpec: Speculative Decoding for Batch Inference of LLM Agents

This work proposes AgentSpec, a speculative decoding algorithm that addresses the limitations of existing methods for LLM agents and incorporates structure-isolated drafting that constrains speculation to semantically coherent segments of the agent workflow, reducing the drafts of irrelevant semantic paths and achievin...

Xin Wang, Zi-Ming Miao, Yi Zhu et al. · 0 citations
#machine learning Preprint Sep 2026

Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: out...

Hao-Yu Zheng, Fang-Cheng Fu, Bin-Hang Yuan et al. · 0 citations
Open access 2026

DABO: Difficulty-Aware Binary Offloading for Collaborative Large-Small Model Inference

DABO is proposed, a calibration-aware binary offloading method for collaborative large–small model inference that maintains competitive end-to-end accuracy while processing an average of 83.72% of requests at the edge.

Chen Zhu, Yi-Ming Su, Chenwenjie Mao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.