2026· International Conference on Conceptual Structures· pp. 472-479· 0 citations· 14 references
Computer Science
TL;DR
It is found that Docker has a negligible impact on performance, while KVM causes a performance drop of 21.22% on average compared to bare metal, and power caps that minimize EDP are different for KVM and Docker/bare metal, which is an important observation for optimization of that essential performance-energy trade-off.
This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preser...
Hanqing Li, Tie-Jun Li, Sheng Ma et al.· ACM Transactions on Design A...· 0 citations
Ghost is an OS-level GPU virtualization layer integrated directly into the open-source GPU driver, using a GPU container abstraction with cgroup -like APIs for compute and memory control and privileged hardware-level scheduling and preemption for dynamic compute resource management.
These findings demonstrate that contemporary generative workloads can be effectively supported by decentralized edge infrastructures, providing practical insights for the design of energy-efficient and heterogeneous local AI systems.
Italo Thiago Felix dos Santos, Felipe Peres De Almeida, Luis Cuevas Rodriguez et al.· Anais do XVIII Simpósio Bras...· 1 citation
This work investigates the feasibility of reproducing benchmarks originally run on datacenter GPUs such as the NVIDIA A100 and RTX 8000 using consumer-grade graphics cards, focusing on the NVIDIA GeForce RTX 3050 and GTX 1060 with CUDA Graphs support. Seven NAS Parallel Benchmarks (BT, LU, SP, EP, IS, MG, and CG) are e...
Leandro L. Retzlaff, Calebe C. Pereira, Helena P. Veltri et al.· Anais do LIII Seminário Inte...· 0 citations
Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders. While Python frameworks have enabled an easy entry point for quantum algorithm design, the low-latency requirements for real-time q...
Joseph K. L. Lee, M. Malekmohammadi, Hong-Sheng Zheng et al.· 0 citations
The high-performance computing industry is moving beyond an era in which each generation of GPU provides uniform performance gains across all applications. The growing importance of AI is driving GPU architecture towards greater specialization, with more silicon devoted to Tensor Cores and reduced-precision arithmetic....
Matthew Tindale, I. Karlin, Tobias Salamon et al.· Inquiry@Queen's Undergraduat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.