The Federated Multi-Armed Bandit (FMAB) framework is proposed to facilitate collaborative model training in cloud-edge environments. Most existing FMAB-based systems assume that participants have personal datasets. This assumption becomes biased in scenarios with limited local data or when the discrepancy between historical data and online data cannot be quantified. Consequently, participants must passively collect data by offering services and gathering feedback from users within their charging areas. In addition, due to the mobility of the users, the number of requests is not stable. In such scenarios, effective model training requires addressing two key challenges. First, the allocation of resources for data collection at the edge is often mismatched with the actual number of service requests, resulting in limited training data and wasted resources. Second, in areas with sparse service requests, the lack of data further delays model adaptation. In this article, networks with these challenges are summarized as the training while collecting data federated bandit (TCF-bandit), and the over-area over-period upper confidence bound (<inline-formula><tex-math notation="LaTeX">$\mathcal {O}^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mi mathvariant="script">O</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="xu-ieq2-3697698.gif"/></alternatives></inline-formula>-UCB) algorithm is proposed to address two challenges. In the <inline-formula><tex-math notation="LaTeX">$\mathcal {O}^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mi mathvariant="script">O</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="xu-ieq3-3697698.gif"/></alternatives></inline-formula>-UCB algorithm, the cloud records request-generation patterns to mitigate resource waste caused by allocation mismatches resulting from unpredictable user mobility. Additionally, a weight-based global confidence radius is computed to assist areas with limited data in quickly identifying their optimal arms. Finally, we prove that the resource allocation waste and regret of the <inline-formula><tex-math notation="LaTeX">$\mathcal {O}^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mi mathvariant="script">O</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="xu-ieq4-3697698.gif"/></alternatives></inline-formula>-UCB algorithm exhibit sub-linear growth. We conduct experiments in different scenarios on the MovieLens, CIFAR-10, and CIFAR-100 datasets to illustrate its superiority over SOTA methods by around 16.2%.
Hang-Fan Li, Yang Xu, Yi-Bin Cai et al.· IEEE Transactions on Mobile...· 0 citations
Retrieval-Augmented Generation (RAG) grounds large language models in external knowledge and has become a key technique for knowledge-intensive tasks. As knowledge bases continue to scale, however, the retrieval stage increasingly dominates end-to-end latency, limiting the responsiveness of RAG systems. In this paper, we identify and empirically validate a previously underexplored property of RAG workloads: strong per-user query locality, where individual users' queries concentrate on a small subset of the knowledge space. Motivated by this observation, we propose Lever, a locality-aware collaborative retrieval framework that exploits query locality to accelerate graph-based RAG retrieval. Lever maintains compact, personalized subgraph indexes on user's local devices as auxiliary structures to guide retrieval toward semantically relevant regions of a global index, enabling more efficient graph traversal without sacrificing coverage. To sustain effectiveness over time, Lever further incorporates adaptive resampling mechanisms that align on-device indexes with evolving query patterns. Extensive experiments on multiple RAG benchmarks demonstrate that Lever significantly reduces retrieval latency and improves throughput while preserving retrieval quality, highlighting query locality as a powerful and complementary lever for scalable RAG retrieval.
Yongheng Deng, Tianyuan Jiang, Zhenya Ma et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.