2026· Global Journal of Pure and Applied Sciences· 0 citations
TL;DR
This study provides a comprehensive analysis of cache memory, including its historical evolution, hierarchical levels (L1–L3), architectural design, and functional significance in modern computing systems, and indicates that cache size alone does not present a statistically significant difference between AMD and Intel processors.
Abstract
Cache memory is a critical determinant of computer system performance, serving as a high-speed intermediary between the processor and the main memory to mitigate the Von Neumann bottleneck. This study provides a comprehensive analysis of cache memory, including its historical evolution, hierarchical levels (L1–L3), architectural design, and functional significance in modern computing systems. The study evaluated state-of-the-art cache architectures, tracking performance through key efficiency metrics including data throughput in Gigabytes per second (GB/s), memory access latency in nanoseconds (ns), and power consumption in picojoules per bit (pJ/bit). While these architectures offer strengths such as reduced latency and improved energy efficiency, they face clear limitations in cost, scalability, and workload dependency. Empirical performance data were compiled from Advanced Micro Devices (AMD) and Intel processors, specifically the AMD Ryzen™ 9 HX PRO 475, 7 PRO 450, 5 PRO 440, Intel® Core™ i9-10850K, i7-1160G7, and i5-1130G7, and analyzed using Analysis of Variance (ANOVA) to compare cache performance. The results indicate that cache size alone does not present a statistically significant difference between AMD and Intel processors. Architectural design and cache management strategies substantially influence the performance outcomes, with Intel processors exhibiting superior cache efficiency under the tested conditions. The findings underscore the pivotal role of cache memory in enhancing processor speed, energy efficiency, and overall system performance, while guiding future innovations in hybrid and adaptive cache architectures.
Cache memory plays a vital role in improving computer system performance by reducing the speed gap between the processor and main memory. This study provides a comparative analysis of various cache memory optimization techniques, including cache replacement policies, mapping methods, multi-level cache architectures, prefetching, and cache partitioning. Using a literature-based approach, the research reviews and evaluates findings from existing studies and scholarly publications. Results indicate that techniques such as Least Recently Used (LRU) policies improve hit rates by 10-25% compared to FIFO, while set-associative mapping reduces miss rates by 15-30% relative to direct mapping. Multi-level cache reduces average memory access latency by up to 50%, and prefetching can boost performance by 20-40% in data-intensive workloads. This study provides a concise overview of cache optimization techniques and their impact on system performance, contributing to a better understanding of cache memory design in modern computer architectures.
Gay Marie P. Farnazo· International journal of res...· 0 citations
This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preserving L2 cache performance.
Hanqing Li, Tiejun Li, Sheng Ma et al.· ACM Transactions on Design A...· 0 citations
Three fundamental design principles are revealed that provide design-space guidance for architects designing the next generation of memory-accelerated LLM systems.
Corey Lammie, Hadjer Benmeziane, W. Simon et al.· 0 citations
It is found that hardware offload is not a universal replacement for CPU compression, and compression should be scheduled dynamically: route blocks by codec, operation, size, and accelerator load; cap per-device sub-mission concurrency; and fall back to software when offload is unsupported or saturated.
Yi Jiang, Antonio Boffa, Hamish Nicholson et al.· 0 citations
MEPOWER is proposed, a flexible, model-based approach to exposing compute/data movement imbalance that characterizes the fine-grained memory behavior of parallel workloads that demonstrates a reduction in EDP on a range of HPC benchmarks with minimal impact on execution time when compared to the standard OS/hardware-managed power control mechanism.
Nanda Velugoti, Joseph Manzano, Andrés Márquez et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.