Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and causes the cache to grow linearly with decoding length. Existing KV-cache compression methods typically control this growth through token eviction, but...
Yi-Ren Zhao, Guang-Hui Song, Tianrui Qin et al.· 0 citations
Agents using the Model Context Protocol (MCP) rely on semantic matching to select tools from third-party servers, exposing a semantic supply-chain risk through attacker-controlled metadata and outputs. We introduce A2M (Attraction-to-Manipulation), a two-stage black-box framework for hijacking MCP agents. The Attractio...
Lai-Zhen Li, Xuan Wang, Pei-Cheng Zhao et al.· 0 citations
Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for ta...
Lai-Zhen Li, Jia-Rui Li, Juanjuan Zhao et al.· 1 citation
CoScale is proposed, a layer-wise dynamic scaling mechanism that selectively expands the parallelism of hot layers onto idle resources reclaimed from underutilized devices, enabling elastic data parallelism without altering model architectures or adding hardware overhead.
Jingfeng Wu, Yiyuan He, Minxian Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.