Skip to content
Open access

Cache memory architecture: A comparative analysis of Intel and AMD cache memory systems

2026 · Global Journal of Pure and Applied Sciences · 0 citations

TL;DR

This study provides a comprehensive analysis of cache memory, including its historical evolution, hierarchical levels (L1–L3), architectural design, and functional significance in modern computing systems, and indicates that cache size alone does not present a statistically significant difference between AMD and Intel processors.

Abstract

Cache memory is a critical determinant of computer system performance, serving as a high-speed intermediary between the processor and the main memory to mitigate the Von Neumann bottleneck. This study provides a comprehensive analysis of cache memory, including its historical evolution, hierarchical levels (L1–L3), architectural design, and functional significance in modern computing systems. The study evaluated state-of-the-art cache architectures, tracking performance through key efficiency metrics including data throughput in Gigabytes per second (GB/s), memory access latency in nanoseconds (ns), and power consumption in picojoules per bit (pJ/bit). While these architectures offer strengths such as reduced latency and improved energy efficiency, they face clear limitations in cost, scalability, and workload dependency. Empirical performance data were compiled from Advanced Micro Devices (AMD) and Intel processors, specifically the AMD Ryzen™ 9 HX PRO 475, 7 PRO 450, 5 PRO 440, Intel® Core™ i9-10850K, i7-1160G7, and i5-1130G7, and analyzed using Analysis of Variance (ANOVA) to compare cache performance. The results indicate that cache size alone does not present a statistically significant difference between AMD and Intel processors. Architectural design and cache management strategies substantially influence the performance outcomes, with Intel processors exhibiting superior cache efficiency under the tested conditions. The findings underscore the pivotal role of cache memory in enhancing processor speed, energy efficiency, and overall system performance, while guiding future innovations in hybrid and adaptive cache architectures.

Read PDF

Similar papers

Review Open access 2026

Performance Analysis of Cache Memory Optimization Techniques in Modern Computer Architectures

Cache memory plays a vital role in improving computer system performance by reducing the speed gap between the processor and main memory. This study provides a comparative analysis of various cache memory optimization techniques, including cache replacement policies, mapping methods, multi-level cache architectures, prefetching, and cache partitioning. Using a literature-based approach, the research reviews and evaluates findings from existing studies and scholarly publications. Results indicate that techniques such as Least Recently Used (LRU) policies improve hit rates by 10-25% compared to FIFO, while set-associative mapping reduces miss rates by 15-30% relative to direct mapping. Multi-level cache reduces average memory access latency by up to 50%, and prefetching can boost performance by 20-40% in data-intensive workloads. This study provides a concise overview of cache optimization techniques and their impact on system performance, contributing to a better understanding of cache memory design in modern computer architectures.

Gay Marie P. Farnazo · 0 citations
Open access Aug 2026

SAI: Virtualizing Shared Memory of GPU for AI workload acceleration

This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preserving L2 cache performance.

Hanqing Li, Tiejun Li, Sheng Ma et al. · 0 citations
Preprint Aug 2026

On Design Principles for Efficient Heterogeneous DRAM-PIM-GPU Systems

Three fundamental design principles are revealed that provide design-space guidance for architects designing the next generation of memory-accelerated LLM systems.

Corey Lammie, Hadjer Benmeziane, W. Simon et al. · 0 citations

The Free Lunch Has a Queue: Characterizing On-chip Compression Accelerators for Analytics

It is found that hardware offload is not a universal replacement for CPU compression, and compression should be scheduled dynamically: route blocks by codec, operation, size, and accelerator load; cap per-device sub-mission concurrency; and fall back to software when offload is unsupported or saturated.

Yi Jiang, Antonio Boffa, Hamish Nicholson et al. · 0 citations
Preprint Aug 2026

MEMPOWER: Efficient Power Management with Fine-grained Memory Analysis and Modeling for HPC Workloads

MEPOWER is proposed, a flexible, model-based approach to exposing compute/data movement imbalance that characterizes the fine-grained memory behavior of parallel workloads that demonstrates a reduction in EDP on a range of HPC benchmarks with minimal impact on execution time when compared to the standard OS/hardware-managed power control mechanism.

Nanda Velugoti, Joseph Manzano, Andrés Márquez et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.