Skip to content
Preprint

STEAM: A Spatio-TEmporal Alignment Mixture-of-Experts Model with Hierarchical Pre-training for EEG Decoding

Aug 2026 · 0 citations · 47 references
Computer Science

TL;DR

STEAM is presented, a hierarchical transfer framework that reconciles general-purpose representation learning with paradigm-specific specialization in EEG foundation models and attains the best average rank among the compared methods at a competitive inference cost measured in FLOPs.

Abstract

Brain-computer interfaces (BCIs) have been widely used in motor rehabilitation, disease diagnosis, and other neural engineering scenarios. However, conventional neural signal decoding algorithms often suffer from limited generalizability and high adaptation costs, motivating recent interest in BCI foundation models. Existing approaches still struggle to jointly achieve general transferability, accurate decoding, and efficient downstream adaptation. We present STEAM, a hierarchical transfer framework that reconciles general-purpose representation learning with paradigm-specific specialization in EEG foundation models. The framework is instantiated as a dual-branch spatio-temporal encoder in which a shared soft mixture-of-experts (SSMoE) module aligns the spatial and temporal branches, allowing complementary representations to exchange information through a compact set of soft slots. Across seven downstream datasets and fourteen evaluation settings, STEAM attains the best average rank among the compared methods at a competitive inference cost measured in FLOPs. Building upon the Stage-I general initialization, the hierarchical pre-training strategy further specializes the model to a target paradigm without retraining from scratch, yielding consistent gains in paradigm-specific decoding accuracy.

View source

Similar papers

Aug 2026

HCFT: A Hierarchical Convolutional Fusion Transformer for Cross-Task EEG Decoding.

Electroencephalography (EEG) decoding remains challenging due to the non-stationary nature of neural signals and the limited generalization of existing models across tasks and subjects. To address this challenge, we propose a lightweight and generalizable decoding framework named Hierarchical Convolutional Fusion Transformer (HCFT), which combines dual-branch convolutional encoders and hierarchical Transformer blocks for multi-scale EEG representation learning. Specifically, the model first captures local temporal and spatiotemporal dynamics through time-domain and time space convolutional branches, and then aligns these features via a cross-attention mechanism that enables interaction between branches at each stage. Subsequently, a hierarchical Transformer fusion structure is employed to encode global dependencies across all feature stages. A task-adaptive stabilization strategy of Dynamic Tanh normalization is introduced to enhance transient feature detection and training stability. Extensive experiments are conducted on two representative cross-task benchmark datasets, BCI Competition IV-2b and CHB-MIT, covering both event-related classification and continuous seizure prediction tasks. Results show that HCFT achieves 80.83% average accuracy and a Cohen's kappa of 0.6165 on BCI IV 2b, as well as 99.10% sensitivity, 0.0236 false positives per hour, and 98.82% specificity on CHB-MIT, consistently outperforming over ten state-of-the-art baseline methods. Ablation studies confirm the effect of each core component of the proposed framework. The model also exhibits strong cross-subject generalization and structural interpretability, offering a scalable and versatile framework for advancing general-purpose neural decoding systems.

Haodong Zhang, Jiapeng Zhu, Yitong Chen et al. · 0 citations
Jul 2026

EEGForceFusion: Joint Tokenised-Continuous Representation Learning for Subject-Independent Grasp Force Decoding

Brain-machine interfaces provide a link between neural activity and external devices, enabling restoration of motor function and advancing human-machine interaction using non-invasive electroencephalography (EEG). However, continuous grasp force decoding remains challenging due to complex temporal dynamics, high inter-subject variability, and limited generalisation of existing approaches. To address this, we propose a hybrid EEG decoding framework that jointly models continuous and tokenised representations, enabling capture of both fine-grained neural structure and long-range temporal dependencies. The proposed approach integrates convolutional-recurrent representation learning, quantisation-based tokenisation, and transformer-based temporal modelling within a unified fusion-based regression architecture. Experimental evaluation on the WAY-EEG-GAL dataset under strict leave-one-subject-out conditions achieves $R^2$ = 0.817 in offline settings and $R^2$ = 0.793 in simulated real-time evaluation, with latency suitable for real-time deployment. These results demonstrate strong cross-subject generalisation and highlight the practicality of hybrid continuous-tokenised representations for real-time EEG-based force decoding in assistive robotics, neuro-rehabilitation, and human-machine interaction.

Sankalp Sunil Turankar, Y. Meena · 0 citations
Review Open access Sep 2026

Cross-variability decoding for motor imagery EEG signals: a comprehensive review

Motor Imagery-based (MI) Electroencephalography (EEG) has emerged as a leading solution in non-invasive Brain-Computer Interface (BCI) systems, leveraging its strong motor intention correlation to enable reliable neural decoding. However, practical implementation of MI confronts three persistent challenges: low signal-to-noise ratio, substantial variability across subjects or over time, and inherent signal nonstationarity. These fundamental limitations continue to hinder the widespread adoption and operational reliability of MI BCI systems. Despite advances in cross-variability decoding methods, there is a lack of systematic syntheses to guide technological evolution in MI BCI. To address these challenges, this review presents a comprehensive taxonomy of MI EEG cross-variability decoding studies from 2020 to 2025, systematically organizing advances in deep learning and transfer learning. We critically evaluate core algorithmic approaches, including Convolutional Neural Networks (CNN), transformers, feature alignment, domain adaptation, and meta-learning. We then explore the underlying mechanisms of these methods and assess their efficacy across key variability paradigms (mainly cross-subject and cross-session scenarios). Finally, we summarize key findings, highlight unresolved challenges, and outline promising future research directions. These advancements hold significant potential to bridge the gap between laboratory-based MI and real-world clinical and consumer applications.

Li-Jun Wang, Yue-Ying Zhou, Peng-Pai Wang et al. · 0 citations
Jul 2026

MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamics Learning

Self-supervised foundation models have recently shown strong potential for electroencephalogram (EEG)-based analysis. However, existing approaches struggle to capture the inherently multi-scale temporal structure of EEG signals, where local neural patterns and long-range dependencies jointly encode task-relevant information. This limitation hampers cross-scale representation learning and generalization across diverse downstream tasks. To address this challenge, we propose MSBraM, a Multi-Scale self-supervised Brain foundation Model designed to learn hierarchical EEG representations. MSBraM follows a two-stage pretraining framework. First, a multi-scale neural tokenizer discretizes raw EEG signals into semantic codes at different temporal resolutions via vector-quantized reconstruction. Second, the model is pretrained to predict masked codes using a curriculum multi-scale masking strategy, progressively integrating fine-grained local patterns with global temporal context. We pretrain MSBraM on over 2,400 hours of EEG data and evaluate it across 10 downstream tasks on 12 public datasets. Extensive experiments show that MSBraM achieves superior performance on other state-of-the-art pretrained models, demonstrating strong generalization and transferability. These results indicate that explicitly modeling multi-scale temporal dynamics is critical for effective EEG foundation models.

Tao Zhou, Jing Han, Lingyu Shu et al. · 0 citations
Open access Sep 2026

SEDAT: a hybrid tokenizer for large EEG models

Objective. The fidelity of neural representations learned by large EEG foundation models depends on how raw brain signals are tokenized. Existing methods suffer from arbitrary temporal boundaries misaligned with neural state transitions, neglecting inter-channel spatial information, and fixed segmentation criteria that fail to generalize across heterogeneous EEG paradigms. Approach. This study proposes the squeeze-and-excitation (SE)-data-adaptive Gaussian average filtering (DAGAF) adaptive tokenizer (SEDAT), a hybrid framework integrating SE-based spatial aggregation, DAGAF-based signal decomposition, instantaneous-frequency-guided adaptive segmentation, and Fourier-domain resampling into a single computationally efficient pipeline. SEDAT is evaluated across 10 heterogeneous EEG datasets spanning motor imagery, mental imagery, P300, slow cortical potentials, sleep staging, epilepsy, emotions recognition and Alzheimer diagnosis paradigms, using four large foundation models: LaBraM, EEGFormer, EEGPT, and NeuroGPT. It is benchmarked against five competitive baselines: fixed-length windowing (FLW), context segmentation, linear predictive coding-based tokenization, TFM-Tokenizer, and source informed segmentation. Main results. SEDAT achieves classification improvements of up to 15.3% over FLW and 1.2%–4.6% over the second-best method, with all comparisons reaching p<0.001 after Benjamini–Hochberg correction. Token quality analysis confirms substantially improved feature separability, with Silhouette scores of 0.81–0.85 versus 0.33–0.48 for rigid baselines. Significance. With O(CN+KNlog⁡N) complexity, SEDAT explores new applications for SE and DAGAF as tokenizers and provides a physiologically grounded and computationally practical tokenization solution for large-scale EEG foundation models.

Muhammad Zulkifal Aziz, Yue Zhuo, Binwen Huang et al. · 0 citations
Open access Aug 2026

FoME: A foundation model for EEG using adaptive temporal-lateral attention scaling.

Electroencephalography (EEG) is a vital tool to measure and record brain activity in neuroscience and clinical applications, yet its potential is constrained by signal heterogeneity, low signal-to-noise ratios, and limited labeled datasets. In this paper, we propose FoME (Foundation Model for EEG), a novel approach using adaptive temporal-lateral attention scaling to address above-mentioned challenges. FoME is pre-trained on a diverse 1.7TB dataset of scalp and intracranial EEG recordings, comprising 745M parameters trained for 1,096k steps. Our model introduces two key innovations: a time-frequency fusion embedding technique and an adaptive temporal-lateral attention scaling (ATLAS) mechanism. These components synergistically capture complex temporal and spectral EEG dynamics, enabling FoME to adapt to varying patterns across diverse data streams and facilitate robust multi-channel modeling. Evaluations across four downstream tasks demonstrate FoME's superior performance in classification and forecasting applications, consistently achieving state-of-the-art results. To conclude, FoME establishes a new paradigm for EEG analysis, offering a versatile foundation that advances brain-computer interfaces, clinical diagnostics, and cognitive research across neuroscience and related fields. Code will be released upon publication.

Enze Shi, Kui Zhao, Qilong Yuan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.