A split-phase heterogeneous deployment strategy is proposed, and key optimization paths, including operator ecosystem completion and deep operator fusion, are identified.
This study presents the first comprehensive, cross-layer measurement study of mobile LLM inference, uniquely spanning five mainstream frameworks and three hardware backends, and identifies a distinct phase split where NPUs excel at compute-bound prefilling, while CPUs outperform all other backends in memory-bound decoding.
Guanyu Cai, Ruiming Tian, Lang Yang et al.· 0 citations
Three fundamental design principles are revealed that provide design-space guidance for architects designing the next generation of memory-accelerated LLM systems.
Corey Lammie, Hadjer Benmeziane, W. Simon et al.· 0 citations
HeteroMosaic first uses a heterogeneous roofline model to identify when combining iGPU and NPU execution is beneficial and decomposes inference into dependency-preserving micro-batches that expose cross-accelerator overlap and applies trace-guided co-optimization of scheduling and device allocation under practical effects such as memory contention, DVFS, device variation, and NPU runtime overheads.
Gregory Jun, Wesley Pang, E. Richter et al.· arXiv.org· 0 citations
Adaptive Sequence Pipeline Parallel Offloading (SPPO) is proposed, a novel framework that optimizes memory and computational resource efficiency for long-sequence LLM training and develops an adaptive pipeline scheduling approach with a heuristic solver and multiplexed sequence partitioning to improve computational resource efficiency.
Qiaoling Chen, Shenggui Li, Wei Gao et al.· International Conference on...· 0 citations
EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.
Jiamin Cao, Qingxu Li, Yaozhong Liu et al.· Conference on Applications,...· 0 citations
Results show that compact SLMs paired with PEFT provide a practical, energy-aware path to personalized on-device deployment, with the optimal method set by the dominant constraint: LoRA+ for energy and QLoRA for memory.
Kuanysh Akhmetzhanov, Jurn-Gyu Park· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.