Skip to content
Preprint

A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

A hybrid CNN-SSM-Attention backbone for 12-lead ECG classification is introduced, which provides strong supervised baselines under a compact parameter budget and improves transfer, particularly in reduced-label settings and under both full fine-tuning and LoRA-based adaptation.

Abstract

Automatic 12-lead electrocardiogram (ECG) classification requires representations that jointly capture local waveform morphology, long-range temporal dynamics, and cross-lead dependencies, yet integrating these properties within a single efficient architecture remains challenging. This paper introduces a hybrid CNN-SSM-Attention backbone for 12-lead ECG classification. A convolutional stem performs early waveform tokenization and temporal reduction, mixed state-space and depthwise-convolutional blocks model temporal dynamics and local morphology, and a late self-attention stage enables global token interaction at reduced resolution. To improve transfer from unlabeled data, we further develop an ECG-oriented Joint-Embedding Predictive Pretraining (JEPA) framework. Unlike ViT-based JEPA methods that mask patch tokens before the encoder, the proposed method samples span masks at the latent temporal resolution and projects them back to the waveform domain, then predicts clean latent targets from a momentum encoder without waveform reconstruction. Experiments on CPSC2018, Chapman-Shaoxing, and PTB-XL, with pretraining on approximately 350K unlabeled CODE-15 recordings, show that the proposed backbone provides strong supervised baselines under a compact parameter budget. JEPA pretraining further improves transfer, particularly in reduced-label settings and under both full fine-tuning and LoRA-based adaptation. Code: https://github.com/yakoubbazi/Hybrid_ECG_Jepa

View source

Similar papers

Open access Sep 2026

A convolutional attention transformer network for ECG beat classification

The findings suggest that the CAT Network can serve as an effective and reproducible framework for AAMI-compliant ECG beat classification, supporting downstream decision support and large-scale screening, while motivating future work on cross-database generalization, computational optimization for edge deployment, and...

Shanmukha Rao Narsupalli, Rajesh Kumar Pullagura, Rajeswara Rao Gangula · 0 citations
#artificial intelligence Preprint Sep 2026

DR-net-Mamba: Selective State-Space Modeling for Long-Range ECG Time-Series Denoising

Electrocardiogram (ECG) recordings are corrupted by non-stationary noise sources that degrade diagnostic reliability, particularly in ambulatory and long-duration recordings. Deep learning denoisers exist, but convolutional architectures are limited by their receptive field, transformer-based models scale quadratically...

Basile Morel, Samuel Ruipérez-Campillo, Andreas P. Streich et al. · 0 citations
Aug 2026

HCFT: A Hierarchical Convolutional Fusion Transformer for Cross-Task EEG Decoding.

A lightweight and generalizable decoding framework named Hierarchical Convolutional Fusion Transformer (HCFT), which combines dual-branch convolutional encoders and hierarchical Transformer blocks for multi-scale EEG representation learning, and exhibits strong cross-subject generalization and structural interpretabili...

Haodong Zhang, Jiapeng Zhu, Yitong Chen et al. · 0 citations
Open access Sep 2026

Higher-dimensional embedding of time-series data for machine learning

Deep learning has revolutionized image analysis, yet most clinical biosignals, especially multi-lead electrocardiograms (ECGs), remain one-dimensional and awkward for modern vision models. We introduce an orthogonal-polynomial imaging (OPI) framework that encodes 12-lead ECGs into a single two-dimensional im...

Karan Singh, Pingal Pratyush Nath, U. Sinha et al. · 0 citations
Open access Aug 2026

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al. · 0 citations
Book Open access Aug 2026

Bridging ECG and PPG: Latent-Space Prediction for Robust Physiological Analysis

This work proposes a multi-modal Joint-Embedding Predictive Architecture for PPG and ECG and shows that the model learns robust cross-modal alignment without relying on contrastive learning, and demonstrates that performance remains robust even with lower-quality signals.

Zhaoliang Chen, Saurabh Kataria, Min-Xiao Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.