Address translation is a major bottleneck in data-intensive workloads. TLB prefetching can hide translation latency, but existing spatial prefetchers struggle with irregular accesses, while temporal prefetchers store address deltas in fixed-capacity hardware that cannot scale with application memory footprints. Our cha...
K. Kanellopoulos, Konstantinos Sgouras, Harsh Songara et al.· 0 citations
Data-intensive applications move large amounts of data from storage to the compute unit, incurring significant data movement overhead. Storage-centric computing reduces this overhead by moving computation near or inside solid-state drives (SSDs). Enabling it requires modifying SSD policies, e.g., address translation an...
Harshita Gupta, Mayank Kabra, Rakesh Nadig et al.· 0 citations
Operating system (OS) code can account for a substantial share of CPU execution time. First, as application logic is offloaded to heterogeneous accelerators (e.g., GPUs), the CPU increasingly acts as an orchestrator, spending cycles in driver calls, data movement, and synchronization rather than in application code. Se...
Vlad-Petru Nitu, Harsh Songara, Konstantinos Sgouras et al.· 0 citations
Vinor is presented, a hardware-OS cooperative memory allocation substrate that combines software flexibility with hardware-class performance and full-system simulation further evaluates the programmable allocation engine and six allocation libraries, showing that Valinor provides hardware-class performance without sacr...
K. Kanellopoulos, Spiros Galanopoulos, Konstantinos Sgouras et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.