End-edge collaborative inference has emerged as an important trend for deploying deep learning applications on resource-constrained end devices. However, the continuous data interaction between end devices and edge server inevitably raises privacy concerns. Fully homomorphic encryption (FHE) provides a privacy-preserving solution by enabling inference directly on encrypted data, but deploying full FHE inference on the edge suffers from high latency. To address these issues, we propose FHE-EESI, an FHE-based end-edge collaborative split inference framework. In FHE-EESI, the end device executes the initial layers on plaintext, encrypts intermediate features using the residue number system Cheon-Kim-Kim-Song (RNS-CKKS) scheme, and transmits them to the edge server for the remaining FHE inference. To construct an FHE-compatible inference network and enable flexible partitionability, we adopt a stage-wise key management strategy and the Chebyshev polynomial approximation. Furthermore, considering dynamic channel conditions and heterogeneous computational capabilities, we design a dueling double deep Q-network (D3QN)-based dynamic split mechanism to adaptively determine the optimal split point. Experimental results on the CIFAR-10 dataset show that FHE-EESI achieves 83.3% inference accuracy and 183.1 s total latency, effectively providing privacy-preserving inference while maintaining inference efficiency.
Haiyue Zhang, Yiming Liu, Jing Jin et al.· IEEE Transactions on Cogniti...· 0 citations
Large vision-language models (LVLMs) have achieved significant progress in video understanding, yet understanding long videos remains challenging due to the large number of visual tokens and limited context windows. Visual sampling provides a practical solution by selecting an informative subset of frames. However, existing methods typically either rely on relevance-aware sampling, leading to redundant frame selection and insufficient temporal coverage, or adopt a fixed sampling strategy regardless of query type. In this paper, we propose VisualRouter, a training-free and plug-and-play framework for query-grounded visual sampling. VisualRouter first classifies each query as either global or local and then applies the corresponding sampling strategy. For global queries, it employs a relevance-coverage hybrid strategy that preserves temporal coverage while retaining query-relevant visual evidence. For local queries, it adopts an event-aware frame selection strategy that performs event partitioning, segment-level frame allocation, and intra-event frame selection, jointly balancing relevance, coverage, and diversity with a limited number of input frames. Experiments show that VisualRouter consistently improves multiple LVLMs over uniform sampling, achieving gains of 5.2%, 7.7%, and 11.6% on Video-MME, LongVideoBench, and MLVU with Qwen2.5-VL-7B, and outperforming existing training-free visual sampling methods under the same setting.
Haiyue Zhang, Yi Bin, Xun Jiang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.