Jul 2026· Signal Processing and Communications Applications Conference· pp. 1-4· 0 citations· 11 references
Abstract
With the widespread adoption of LLM-based chatbots, cloud-dependent solutions have come to dominate the market. However, open-source pre-trained LLMs are enabling implementation of local solutions. Achieving competitive performance locally requires the ability to run high-parameter models. Here, the primary bottleneck is GPU VRAM capacity, which limits model parameter size. Furthermore, the efficiency of inference optimizations such as kv caching depends directly on the amount of available VRAM remaining after the model is loaded. In such a scenario with hardware constraints, we conduct user tests by applying kv cache quantization. As a result, we identify distinct performance trends in critical metrics such as Time-to-First-Token (TTFT) and Total Generation Time. Additionally, we evaluate model accuracy results using the LLM-as-a-judge paradigm.
Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them across the edge continuum requires balancing quality, latency, model footprint, and energy. This paper presents a controlled measurement study of self-hosted LLM inference across edge and near-edge deployment n...
Maysam Khatib, Moysis Symeonides, Demetris Trihinas et al.· 0 citations
CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.
Zejian Wang, Nan Lin, Zinuo Cai et al.· ACM Transactions on Architec...· 0 citations
DynamoServe is presented, a multi-tenant LLM serving framework that addresses challenges through three key innovations: leveraging stranded GPU memory to offload model weights and KV caches, mitigating resource fragmentation in multi-workload environments, and improving memory locality through coordinated data placemen...
Diman Zad Tootaghaj, Khaled Diab, Bob Lantz et al.· Conference on Applications,...· 0 citations
EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.
Jia-Min Cao, Qingxu Li, Yaozhong Liu et al.· Conference on Applications,...· 0 citations
SLIM (Saturation-Aware Lightweight Performance Model), a semi-analytical model that predicts LLM inference throughput and latency from analytical formulations of Transformer computation and memory traffic, is introduced, which outperforms representative performance-modeling baselines while successfully generalizing to...
Pol G. Recasens, F. Agulló, Yue Zhu et al.· arXiv.org· 0 citations
OptiFlow is presented, among the first LLM-driven frameworks for automated design of high-performance collective communication algorithms, with key insight is a two-layer decomposition: the LLM generates compact data-movement intent expressed in a domain-specific language, while deterministic scheduling algorithms comp...
Fei Long, Ziyue Yang, Kaihui Gao et al.· Asia-Pacific Workshop on Net...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.