Jul 2026· Annual International Computer Software and Applications Conference· pp. 2871-2876· 0 citations· 32 references
Computer Science
Abstract
Coordinating heterogeneous LLM agents under congested online settings is difficult because bursty task arrivals and limited per-agent capacity may induce hotspot overload and severe tail waiting time. This paper proposes an online scheduling method based on candidate-set contraction before assignment. Specifically, tasks are first routed through subscription matching to identify a task-relevant candidate pool, after which layered gating is applied to enforce capability feasibility, historical quality, and real-time load constraints. The remaining candidates are then ranked using a composite score that balances competence and load, with stable tie-breaking introduced to reduce assignment fluctuations under contention. We evaluate the method under a reproducible protocol with both regular and congested regimes. Across benchmarks covering code generation, arithmetic reasoning, and preference-based evaluation, the proposed approach preserves competitive task performance while reducing both mean and 95th-percentile waiting time in congested settings relative to representative linear, flat, and hierarchical baselines. The findings suggest that candidate contraction is a useful strategy for achieving more stable coordination in heterogeneous LLM-agent systems.
Dynamic multitask allocation in heterogeneous environments must reconcile global throughput with the strict protection of urgent tasks under online arrivals, coupled resources, and tight time window constraints. Existing heuristic, learning based, and hybrid schedulers typically address isolated problem facets, degradi...
Ming Lei, You-Chen Fan, Xi Xiao et al.· Journal of King Saud Univers...· 0 citations
This work study deadline-aware, mixed-criticality scheduling on heterogeneous MEC servers, where time-critical (TC) tasks must be protected at a controlled cost to best-effort traffic, and asks whether a multi-agent LLM control layer improves on a strong heuristic.
This paper investigates centralized scheduling of mobile service agents under staggered multi-wave task arrivals, limited service capacity, service-time windows, and cross-wave capacity reservation requirements. A spatiotemporal candidate-arc representation is developed to discretize the continuous scheduling process,...
Miao Shen, Chuan-Fu Guo, Peng Wang et al.· 2026 12th International Conf...· 0 citations
Multi-agent applications increasingly rely on shared large language model backends in the public cloud, where bursty workloads cause requests from different agents to contend for the same LLM instances, leading to long queues, memory imbalance, and severe tail-latency inflation. Existing approaches typically prioritize...
Ali Zafar Sadiq, Hai-Ying Shen· International Conference on...· 0 citations
Agentic LLM workflows issue many dependent calls with unpredictable resource demand, causing queue buildup and latency degradation on shared serving backends when left unmanaged. In this paper, we propose CALM-MAS, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an ela...
Mouheb Ben Nasr, Muhammad Bilal, Alessandro Cornacchia et al.· Proceedings of the 17th ACM...· 0 citations
This work proposes MARA, which predicts future loss trajectories with conditional flow matching and coordinates compute nodes through a cooperative multi-agent autoregressive policy and reduces remaining-resource prediction error relative to weighted least squares.
Han-Ye Zhao, Mu-Ning Wen, Yong Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.