Skip to content
Book Open access

MVP: A Mobile 3D-Stacked VLM Accelerator for Efficient Video Understanding by Leveraging Dynamic Sparse Attention Patterns

Aug 2026 · International Symposium on Low Power Electronics and Design · 0 citations · 15 references
Computer Science

Abstract

The emergence of Vision-Language Models (VLMs) has enabled multimodal reasoning, e.g. video understanding, yet their extension to long-context inference remains bottlenecked by the “token explosion”. This surge in sequence length leads to prohibitive attention computation overhead and memory-bound KV cache access. While 3D-stacked logic-to-DRAM architectures offer high-bandwidth Processing Near Memory (PNM) capability, their distributed memory banks face severe workload imbalance due to the unique spatiotemporal sparsity patterns in video understanding tasks. In this paper, we introduce MVP, a 3D-stacked VLM accelerator featuring context-aware sparse attention (CASA) and online workload-aware hybrid parallelism scheduling. By leveraging dynamic sparse attention patterns, our design prunes redundant attention computation FLOPS and adaptively balances computation across hybrid bonding (HB) based many-core NoC architecture. Experimental results on VQA tasks demonstrate that our architecture achieves a 5.97× speedup and a 5.42× energy-efficiency improvement over the RTX 4080 GPU deployment, effectively mitigating the bottlenecks of VLM inference for video-understanding on mobile devices.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.