Skip to content
Open access

ROI-Guided Echocardiographic Image Analysis and Case Aggregation for Differentiating Atrial and Ventricular Septal Defects

2026 · IEEE Access · Vol 14, pp. 130235-130247 · 0 citations · 35 references

TL;DR

This study implemented strict subject-disjoint partitioning to eliminate data leakage, and simultaneously introduced cross-frame case aggregation to emulate the multi-frame visual synthesis process of expert echocardiographers, suggesting that the proposed workflow has the potential to serve as an adjunctive tool for septal defect screening.

Abstract

Echocardiography is the primary imaging modality for congenital heart disease (CHD) assessment. However, the real-world clinical application of artificial intelligence in this domain is often hindered by hidden data leakage from homologous frames and the high diagnostic variance of single-frame static inference. To bridge the gap between idealized model evaluation and clinical reality, this study proposes a standardized, region-of-interest (ROI)-guided deep learning workflow for differentiating atrial septal defect (ASD), ventricular septal defect (VSD), and normal cases. Specifically, utilizing a retrospective cohort of 1,987 patients (11,638 images) from Tengzhou Central People’s Hospital, we implemented strict subject-disjoint partitioning to eliminate data leakage, and simultaneously introduced cross-frame case aggregation to emulate the multi-frame visual synthesis process of expert echocardiographers. Among the evaluated model configurations, fully fine-tuned EfficientNet-B0 achieved the highest internal held-out test performance, with an accuracy of 0.9950 (95% CI, 0.9875–1.0000), a macro-F1 of 0.9941, and a macro-AUROC of 0.9997; 2 of 401 test patients were misclassified. Five-fold patient-level cross-validation yielded an accuracy of $0.9945 \pm 0.0040$ , further supporting the stability of the model across patient partitions. These findings suggest that the proposed workflow has the potential to serve as an adjunctive tool for septal defect screening.

Read PDF

Similar papers

Aug 2026

Structure-Aware Deep Learning for Pediatric Echocardiographic Standard-View Classification in Ultrasonic Imaging.

Accurate standard-view classification is essential for pediatric echocardiographic image analysis and downstream automated interpretation. This task remains challenging because discriminative view information is often encoded in subtle chamber configurations, outflow-tract morphology, and weak anatomical boundaries, whereas conventional classifiers may underuse shallow and intermediate representations that preserve spatial structure. We propose PVTv2-ASEF, a structure-aware framework for 4-class pediatric echocardiographic standard-view classification. The framework introduces an Adaptive Structural Enhancement Module that performs residual input-side conditioning through channel recalibration, local convolutional mixing, multi-scale structural modeling, and input-dependent branch weighting. It further employs a Dual Auxiliary Fusion Head to transform Stage 2 and Stage 3 representations into class-level evidence and fuse them with the final-stage logits during inference. PVTv2-ASEF was evaluated on a private pediatric ventricular septal defect echocardiography dataset comprising 4 standard views under repeated patient-disjoint evaluation, with macro-F1 used as the primary class-balanced metric. Compared with PVTv2-B2, the proposed framework improved macro-F1 by 0.093 on the private dataset and by 0.023 in cross-task evaluation on FETAL_PLANES_DB. These results support the utility of coordinated input-side structural enhancement and intermediate logit fusion for ultrasound view and plane classification.

Yanfeng Liu, Haibin Sun, Hai-Song Huang et al. · 0 citations
Open access Jul 2026

Comprehensive Echocardiography Interpretation Using Video and Multiview Vision-Language AI.

A multiview video-language framework improved report retrieval compared with conventional image-based approaches and support the utility of video-based, multiview representation learning for echocardiographic report retrieval.

R. Takizawa, Chiemi Yamazaki, S. Kodera et al. · 0 citations
Jul 2026

Domain Shift in Echocardiography: Interpretable Quantification and Prediction of Cross-Dataset Left Ventricular Segmentation

Cross-dataset generalisation remains a major barrier to clinical deployment of echocardiographic left ventricular segmentation, yet the sources of this shift are rarely disentangled. We examined whether transfer degradation could be estimated before deployment using handcrafted ultrasound descriptors, VAE latent features, and segmentation-derived latent features across six echocardiographic datasets. Geometry-aware preprocessing substantially improved several poor transfer cases, suggesting that much of the apparent domain shift reflects field-of-view and framing inconsistencies rather than intrinsic acoustic differences alone. Intensity z-normalisation changed dataset separability by less than 0.005, indicating that brightness and contrast are not the dominant shift axis. Absolute Dice drop on held-out source-target pairs was predicted with an R-squared value of 0.612, an MAE of 0.082, and a Spearman rho of 0.681. The variant without LV and fan-shaped features retained approximately 70% of this explanatory power, supporting mask-free transfer-risk monitoring. The most informative discrepancy measure depended on the representation, with CMD strongest in z-normalised handcrafted features, with an absolute r of approximately 0.86 and an R-squared value of approximately 0.70; log-Wasserstein strongest in VAE space, with an r of approximately -0.90 and an R-squared value of approximately 0.81; and log-MMD strongest in LV-segmentation latent features, with an r of approximately -0.92 and an R-squared value of approximately 0.84. Apparent vendor effects were largely dataset-confounded. Echocardiographic domain shift is therefore structured and measurable, and its impact on segmentation can be partly reduced through geometry-aware preprocessing and anticipated using representation-specific transfer-risk estimation.

Soroush Elyasi, Nasim Dadashi Serej, Julie Wall et al. · 0 citations
Aug 2026

Beyond Doppler: Scalable AI Detection of LVOT Obstruction in HCM.

BACKGROUND Accurate assessment of left ventricular outflow tract (LVOT) gradients is critical for hypertrophic cardiomyopathy management, yet Doppler-based measurements are technically demanding and require expertise. The objective of this work was to develop a multi-view deep learning model capable of classifying LVOT obstruction (>20 mm Hg) using routine 2-dimensional echocardiographic windows without reliance on Doppler imaging. METHODS We trained and externally validated a cross-attention-based video-to-video fusion framework that integrated EchoPrime-derived video representations from 3 standard transthoracic echocardiographic views to classify LVOT gradients. RESULTS Training was performed on a derivation cohort (N=1833) from a tertiary care system in the United States, with model performance evaluated on an internally held-out test set (N=275) and a Korean external validation cohort (N=46). Single-view baselines showed limited discrimination (external area under the receiver operating curves, 0.47-0.70). Conversely, the domain-specific foundational model (EchoPrime) achieved superior single-view performance (area under the receiver operating curves, 0.75-0.80 internal; 0.79-0.83 external), highlighting the importance of echo-specific pretraining and temporal modeling. The proposed multi-view fusion further enhanced predictive performance, with the late fusion model reaching an area under the receiver operating curve of 0.84 on the external cohort with significant population-shift. CONCLUSIONS These results suggest LVOT physiology is encoded in routine 2-dimensional imaging and can be leveraged for clinically relevant gradient classification without Doppler input. The proposed artificial intelligence-guided strategy demonstrates substantial cost savings compared with the screen-all approach. By integrating complementary spatial-temporal information across multiple views, our approach generalizes robustly across populations and may enable real-time decision support, extend LVOT assessment to portable or resource-limited settings, and complement Doppler-based evaluation for longitudinal hypertrophic cardiomyopathy management.

O. Crystal, J. Farina, I. Scalia et al. · 0 citations
Open access Jul 2026

BackMix-Enhanced Semi-Supervised Learning for Automated Detection of Aortic Stenosis from Transthoracic Echocardiographic Images

The proposed anatomically guided BackMix augmentation combined with semi-supervised ensemble learning can improve classification accuracy, robustness, and interpretability in echocardiographic analysis under limited annotation conditions, offering a promising approach for automated AS assessment across independent clinical datasets.

Fatima Ezzahra Elkouahy, Badreddine Labakoum, H. Ouahid et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.