Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs...
Zhe-Wei Fang, Yu-Xin Zhang, Zhen-Wei Shao et al.· 0 citations
On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and cost. Although KV caching is widely used to reduce latency in long-context inference, existing designs were primarily opt...
Zheng-Xiang Huang, Sheng-Heng Chen, Chao-Yue Niu et al.· 0 citations
RecGPT-Mobile-V2 is introduced, an end-to-end framework that treats intent quality and execution efficiency as coupled objectives within a staged design and helps retain decision-relevant evidence and allocate additional computation only when it is likely to improve the predicted Query.
Lingqin Zhang, Bin Zhang, Wei-Peng Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.