APAC: Adaptive Prompting for Agent-to-Agent Communication at the Edge
Abstract
The rapid deployment of Large Language Model (LLM)-based agents in wireless networks has introduced severe communication bottlenecks due to the continuous exchange of massive prompt sequences. Existing prompt compression methods either cause semantic degradation or introduce prohibitive computational delays, and largely fail to jointly optimize network variables and compression strategies for ultra-low latency multi-agent collaboration. To address these challenges, we propose the Adaptive Prompting Agent-to-Agent Communication (APAC) framework, a dynamic prompt compression mechanism tailored for edge networks. By modeling the joint compute-communicate process as a highly constrained optimization problem, APAC dynamically toggles between lightweight extractive token pruning and high-fidelity generative semantic compression, while continuously adapting the compression ratio based on real-time bandwidth, prompt lengths, and hardware pipelining capabilities. We develop a custom Proximal Policy Optimization (PPO) algorithm tailored for hybrid action spaces to seamlessly balance local computational delay, transmission overhead, and task-conditioned semantic utility. Extensive evaluations demonstrate that APAC achieves a Pareto optimal trade-off between latency and semantic preservation, exhibiting remarkable resilience and strictly circumventing catastrophic latency violations under network congestion and scaling prompt lengths.