From Cray to Exascale: A Critical Survey of Supercomputer Architecture Evolution, Heterogeneity, and Interconnect Bottlenecks
The evolution of supercomputer architecture has undergone several transformative phases since the 1960s, yet existing surveys have not adequately captured the engineering trade-offs that defined each generation. This paper presents a critical survey of supercomputer architecture from early vector machines to contemporary exascale systems, with particular emphasis on three persistent challenges: memory hierarchy design, interconnection network scalability, and thermal management. Drawing on foundational work by Cray and subsequent massively parallel systems, we examine how the transition from shared memory to distributed architectures enabled processor counts to grow from dozens to millions. The paper analyses recent exascale systems including Frontier, Fugaku, and LUMI, alongside emerging accelerator technologies such as GPU-based nodes, wafer-scale integration, and specialized interconnects including Slingshot, InfiniBand, and Omni-Path. Our survey identifies that while peak performance has grown exponentially, the gap between theoretical and sustained performance remains significant, largely due to interconnect latency and memory bandwidth limitations. We further evaluate contemporary cooling approaches from immersion to direct to-chip liquid systems, noting that power density has emerged as the primary constraint on further scaling. Building upon earlier work by Oraye and Anireh (2022), this survey provides updated taxonomies and performance metrics that inform future exascale and post-exascale designs.