Aug 2026· International Symposium on Low Power Electronics and Design· pp. 1-7· 0 citations· 30 references
Computer Science
TL;DR
Software simulations and hardware evaluation show that the proposed operation fusion for the attention training dataflow using local safe softmax and model-independent tiling reduces off-chip access, on-chip memory usage, FLOPs, training runtime, and energy cost compared to conventional approaches, confirming its suitability for efficiently training attention in resource-restricted hardware.
Abstract
The demand for efficient training of Transformer language models is rapidly increasing, yet the resource-intensive attention mechanism severely bottlenecks the process. Previous attempts to fuse attention operations managed to reduce off-chip memory access, thereby improving memory bandwidth. Unfortunately, these techniques are impractical in restricted environments, as they exhibit resource utilization that scales with model dimensions and do not consider the entire training dataflow. To address this, we propose operation fusion for the attention training dataflow using local safe softmax and model-independent tiling. By taking both the forward and the backward passes into account, local safe softmax reduces I/O access under small-cache conditions. Furthermore, model-independent tiling ensures that the required on-chip memory footprint remains independent of model dimensions, enabling scalability across diverse models. Software simulations and hardware evaluation show that our method reduces off-chip access, on-chip memory usage, FLOPs, training runtime, and energy cost compared to conventional approaches, confirming its suitability for efficiently training attention in resource-restricted hardware.
This work proposes a BP-free algorithm, called ZeroLock, that decouples the model updates into independent chunk updates by local objective construction and provides the first theoretical framework for such local objective construction-based approaches under general model chunk division by mapping local objectives to t...
Wen-Tao Dai, Xuan-Ran Li, Yu-Xiang Zhang et al.· 0 citations
This paper presents FlashAttention-V, a blocked FlashAttention for scalable vector architectures that adapts efficiently from short to very long vectors by exploiting parallelism across attention heads, inter-head packing to enable efficient utilization of vector lengths beyond the head dimension, and improving vector...
Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...
Xiaoyang Sun, Jie Xu, Zheng Wang· IEEE Transactions on Paralle...· 0 citations
Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.
Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al.· 3 citations
While Parameter-Efficient Fine-Tuning (PEFT) has substantially reduced the hardware cost of adapting Large Language Models (LLMs) by decreasing the number of trainable parameters, recent studies have sought to further improve PEFT through parameter sharing. However, these approaches either employ uniform parameter shar...
Mohammad Aref Jafari-Raddani, M. M. Kafshdooz· 0 citations
A3D-MoE addresses large language models' challenges with 3-D heterogeneous integration to improve memory bandwidth and reduce NoC overhead/energy, and a hardware resource-aware operation fusion scheduler that fuses attention/MoE operations to boost performance.
Wei-Hsing Huang, Janak Sharda, Cheng-Jhih Shih et al.· IEEE Journal on Exploratory...· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.