Skip to content
Book Open access

Fusing the Attention Training Dataflow with Local Safe Softmax and Model-Independent Tiling

Aug 2026 · International Symposium on Low Power Electronics and Design · pp. 1-7 · 0 citations · 30 references
Computer Science

TL;DR

Software simulations and hardware evaluation show that the proposed operation fusion for the attention training dataflow using local safe softmax and model-independent tiling reduces off-chip access, on-chip memory usage, FLOPs, training runtime, and energy cost compared to conventional approaches, confirming its suitability for efficiently training attention in resource-restricted hardware.

Abstract

The demand for efficient training of Transformer language models is rapidly increasing, yet the resource-intensive attention mechanism severely bottlenecks the process. Previous attempts to fuse attention operations managed to reduce off-chip memory access, thereby improving memory bandwidth. Unfortunately, these techniques are impractical in restricted environments, as they exhibit resource utilization that scales with model dimensions and do not consider the entire training dataflow. To address this, we propose operation fusion for the attention training dataflow using local safe softmax and model-independent tiling. By taking both the forward and the backward passes into account, local safe softmax reduces I/O access under small-cache conditions. Furthermore, model-independent tiling ensures that the required on-chip memory footprint remains independent of model dimensions, enabling scalability across diverse models. Software simulations and hardware evaluation show that our method reduces off-chip access, on-chip memory usage, FLOPs, training runtime, and energy cost compared to conventional approaches, confirming its suitability for efficiently training attention in resource-restricted hardware.

Read PDF

Similar papers

Preprint Aug 2026

ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

This work proposes a BP-free algorithm, called ZeroLock, that decouples the model updates into independent chunk updates by local objective construction and provides the first theoretical framework for such local objective construction-based approaches under general model chunk division by mapping local objectives to t...

Wen-Tao Dai, Xuan-Ran Li, Yu-Xiang Zhang et al. · 0 citations
#small language model Preprint Aug 2026

FlashAttention for Scalable Vector Architectures

This paper presents FlashAttention-V, a blocked FlashAttention for scalable vector architectures that adapts efficiently from short to very long vectors by exploiting parallelism across attention heads, inter-head packing to enable efficient utilization of vector lengths beyond the head dimension, and improving vector...

Sonia Rani Gupta, Nikela Papadopoulou, Miquel Pericàs · 0 citations
Oct 2026

SynergyScale: Optimizing Offloading and Task Partitioning for Efficient Model Training

Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...

Xiaoyang Sun, Jie Xu, Zheng Wang · 0 citations
#artificial intelligence Preprint Aug 2026

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.

Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al. · 3 citations
Preprint Aug 2026

SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning

While Parameter-Efficient Fine-Tuning (PEFT) has substantially reduced the hardware cost of adapting Large Language Models (LLMs) by decreasing the number of trainable parameters, recent studies have sought to further improve PEFT through parameter sharing. However, these approaches either employ uniform parameter shar...

Mohammad Aref Jafari-Raddani, M. M. Kafshdooz · 0 citations
Open access Jul 2025

A3D-MoE: Acceleration of Large Language Models With Mixture of Experts via 3-D Heterogeneous Integration

A3D-MoE addresses large language models' challenges with 3-D heterogeneous integration to improve memory bandwidth and reduce NoC overhead/energy, and a hardware resource-aware operation fusion scheduler that fuses attention/MoE operations to boost performance.

Wei-Hsing Huang, Janak Sharda, Cheng-Jhih Shih et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.