Application-Driven System Technology Co-Optimization for 2.5D Edge AI Platforms
Abstract
Edge AI systems face stringent run-time constraints and tight energy budgets, demanding comprehensive optimization of computing platforms. Nonetheless, existing approaches often focus only on the co-design of hardware and software, but rarely link algorithmic opportunities with system- and technology-level optimizations. Hence, avenues exploiting dynamic machine-learning behaviors on emerging disaggregated System-in-Package platforms remain largely unexplored. To address this challenge, we introduce a cross-level Application-System-Technology Co-Optimization (ASTCO) methodology for dynamic ML encapsulating monolithic and 2.5D chiplet-based architectures. ASTCO explores algorithmic adaptivity, heterogeneous chiplet mapping, and technology scaling within a unified framework. It evaluates integration strategies with or without interposer connectivity and heterogeneous technology nodes, capturing energy-efficiency versus Quality-of-Service (QoS) trade-offs under real-time constraints. We evaluate ASTCO on transformer and CNN models with Early Exits (EEs), representative of dynamic edge ML, performing automated seizure detection. Compared to a monolithic baseline integrating the EE network on a single chip, ASTCO identifies energy-efficient chiplet configurations through cross-level exploration. Results highlight up to 2.2× energy reduction for a 6-layer transformer model, while maintaining real-time guarantees and with very limited QoS impact due to EEs.