Skip to content
Conference

Congestion-Aware Scheduling for Heterogeneous LLM-Agent Teams

Jul 2026 · Annual International Computer Software and Applications Conference · pp. 2871-2876 · 0 citations · 32 references
Computer Science

Abstract

Coordinating heterogeneous LLM agents under congested online settings is difficult because bursty task arrivals and limited per-agent capacity may induce hotspot overload and severe tail waiting time. This paper proposes an online scheduling method based on candidate-set contraction before assignment. Specifically, tasks are first routed through subscription matching to identify a task-relevant candidate pool, after which layered gating is applied to enforce capability feasibility, historical quality, and real-time load constraints. The remaining candidates are then ranked using a composite score that balances competence and load, with stable tie-breaking introduced to reduce assignment fluctuations under contention. We evaluate the method under a reproducible protocol with both regular and congested regimes. Across benchmarks covering code generation, arithmetic reasoning, and preference-based evaluation, the proposed approach preserves competitive task performance while reducing both mean and 95th-percentile waiting time in congested settings relative to representative linear, flat, and hierarchical baselines. The findings suggest that candidate contraction is a useful strategy for achieving more stable coordination in heterogeneous LLM-agent systems.

View source

Similar papers

Open access Aug 2026

AMCATS: An adaptive rule-guided framework for dynamic multi-task allocation in heterogeneous satellite systems

Dynamic multitask allocation in heterogeneous environments must reconcile global throughput with the strict protection of urgent tasks under online arrivals, coupled resources, and tight time window constraints. Existing heuristic, learning based, and hybrid schedulers typically address isolated problem facets, degradi...

Ming Lei, You-Chen Fan, Xi Xiao et al. · 0 citations
Conference Aug 2026

Discrete Modeling and Combinatorial Optimization for Task Scheduling

This paper investigates centralized scheduling of mobile service agents under staggered multi-wave task arrivals, limited service capacity, service-time windows, and cross-wave capacity reservation requirements. A spatiotemporal candidate-arc representation is developed to discretize the continuous scheduling process,...

Miao Shen, Chuan-Fu Guo, Peng Wang et al. · 0 citations
Conference Jul 2026

FlowGuard: Slack-Aware Overload Control for Multi-Agent LLM Serving

Multi-agent applications increasingly rely on shared large language model backends in the public cloud, where bursty workloads cause requests from different agents to contend for the same LLM instances, leading to long queues, memory imbalance, and severe tail-latency inflation. Existing approaches typically prioritize...

Ali Zafar Sadiq, Hai-Ying Shen · 0 citations
Book Open access Sep 2026

Congestion-Aware Serving of Agentic LLM Applications

Agentic LLM workflows issue many dependent calls with unpredictable resource demand, causing queue buildup and latency degradation on shared serving backends when left unmanaged. In this paper, we propose CALM-MAS, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an ela...

Mouheb Ben Nasr, Muhammad Bilal, Alessandro Cornacchia et al. · 0 citations
Preprint Aug 2026

MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning

This work proposes MARA, which predicts future loss trajectories with conditional flow matching and coordinates compute nodes through a cooperative multi-agent autoregressive policy and reduces remaining-resource prediction error relative to weighted least squares.

Han-Ye Zhao, Mu-Ning Wen, Yong Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.