This study implemented strict subject-disjoint partitioning to eliminate data leakage, and simultaneously introduced cross-frame case aggregation to emulate the multi-frame visual synthesis process of expert echocardiographers, suggesting that the proposed workflow has the potential to serve as an adjunctive tool for septal defect screening.
Abstract
Echocardiography is the primary imaging modality for congenital heart disease (CHD) assessment. However, the real-world clinical application of artificial intelligence in this domain is often hindered by hidden data leakage from homologous frames and the high diagnostic variance of single-frame static inference. To bridge the gap between idealized model evaluation and clinical reality, this study proposes a standardized, region-of-interest (ROI)-guided deep learning workflow for differentiating atrial septal defect (ASD), ventricular septal defect (VSD), and normal cases. Specifically, utilizing a retrospective cohort of 1,987 patients (11,638 images) from Tengzhou Central People’s Hospital, we implemented strict subject-disjoint partitioning to eliminate data leakage, and simultaneously introduced cross-frame case aggregation to emulate the multi-frame visual synthesis process of expert echocardiographers. Among the evaluated model configurations, fully fine-tuned EfficientNet-B0 achieved the highest internal held-out test performance, with an accuracy of 0.9950 (95% CI, 0.9875–1.0000), a macro-F1 of 0.9941, and a macro-AUROC of 0.9997; 2 of 401 test patients were misclassified. Five-fold patient-level cross-validation yielded an accuracy of $0.9945 \pm 0.0040$ , further supporting the stability of the model across patient partitions. These findings suggest that the proposed workflow has the potential to serve as an adjunctive tool for septal defect screening.
MitralVision reliably distinguishes clinically significant MR using single-view B-mode echocardiography without Doppler input for model inference and may support more standardized MR screening.
R. Sandler, J. Sokol, S.G. Pawar et al.· Journal of the American Soci...· 0 citations
Accurate standard-view classification is essential for pediatric echocardiographic image analysis and downstream automated interpretation. This task remains challenging because discriminative view information is often encoded in subtle chamber configurations, outflow-tract morphology, and weak anatomical boundaries, whereas conventional classifiers may underuse shallow and intermediate representations that preserve spatial structure. We propose PVTv2-ASEF, a structure-aware framework for 4-class pediatric echocardiographic standard-view classification. The framework introduces an Adaptive Structural Enhancement Module that performs residual input-side conditioning through channel recalibration, local convolutional mixing, multi-scale structural modeling, and input-dependent branch weighting. It further employs a Dual Auxiliary Fusion Head to transform Stage 2 and Stage 3 representations into class-level evidence and fuse them with the final-stage logits during inference. PVTv2-ASEF was evaluated on a private pediatric ventricular septal defect echocardiography dataset comprising 4 standard views under repeated patient-disjoint evaluation, with macro-F1 used as the primary class-balanced metric. Compared with PVTv2-B2, the proposed framework improved macro-F1 by 0.093 on the private dataset and by 0.023 in cross-task evaluation on FETAL_PLANES_DB. These results support the utility of coordinated input-side structural enhancement and intermediate logit fusion for ultrasound view and plane classification.
A multiview video-language framework improved report retrieval compared with conventional image-based approaches and support the utility of video-based, multiview representation learning for echocardiographic report retrieval.
R. Takizawa, Chiemi Yamazaki, S. Kodera et al.· JACC: Asia· 0 citations
Cross-dataset generalisation remains a major barrier to clinical deployment of echocardiographic left ventricular segmentation, yet the sources of this shift are rarely disentangled. We examined whether transfer degradation could be estimated before deployment using handcrafted ultrasound descriptors, VAE latent features, and segmentation-derived latent features across six echocardiographic datasets. Geometry-aware preprocessing substantially improved several poor transfer cases, suggesting that much of the apparent domain shift reflects field-of-view and framing inconsistencies rather than intrinsic acoustic differences alone. Intensity z-normalisation changed dataset separability by less than 0.005, indicating that brightness and contrast are not the dominant shift axis. Absolute Dice drop on held-out source-target pairs was predicted with an R-squared value of 0.612, an MAE of 0.082, and a Spearman rho of 0.681. The variant without LV and fan-shaped features retained approximately 70% of this explanatory power, supporting mask-free transfer-risk monitoring. The most informative discrepancy measure depended on the representation, with CMD strongest in z-normalised handcrafted features, with an absolute r of approximately 0.86 and an R-squared value of approximately 0.70; log-Wasserstein strongest in VAE space, with an r of approximately -0.90 and an R-squared value of approximately 0.81; and log-MMD strongest in LV-segmentation latent features, with an r of approximately -0.92 and an R-squared value of approximately 0.84. Apparent vendor effects were largely dataset-confounded. Echocardiographic domain shift is therefore structured and measurable, and its impact on segmentation can be partly reduced through geometry-aware preprocessing and anticipated using representation-specific transfer-risk estimation.
BACKGROUND
Accurate assessment of left ventricular outflow tract (LVOT) gradients is critical for hypertrophic cardiomyopathy management, yet Doppler-based measurements are technically demanding and require expertise. The objective of this work was to develop a multi-view deep learning model capable of classifying LVOT obstruction (>20 mm Hg) using routine 2-dimensional echocardiographic windows without reliance on Doppler imaging.
METHODS
We trained and externally validated a cross-attention-based video-to-video fusion framework that integrated EchoPrime-derived video representations from 3 standard transthoracic echocardiographic views to classify LVOT gradients.
RESULTS
Training was performed on a derivation cohort (N=1833) from a tertiary care system in the United States, with model performance evaluated on an internally held-out test set (N=275) and a Korean external validation cohort (N=46). Single-view baselines showed limited discrimination (external area under the receiver operating curves, 0.47-0.70). Conversely, the domain-specific foundational model (EchoPrime) achieved superior single-view performance (area under the receiver operating curves, 0.75-0.80 internal; 0.79-0.83 external), highlighting the importance of echo-specific pretraining and temporal modeling. The proposed multi-view fusion further enhanced predictive performance, with the late fusion model reaching an area under the receiver operating curve of 0.84 on the external cohort with significant population-shift.
CONCLUSIONS
These results suggest LVOT physiology is encoded in routine 2-dimensional imaging and can be leveraged for clinically relevant gradient classification without Doppler input. The proposed artificial intelligence-guided strategy demonstrates substantial cost savings compared with the screen-all approach. By integrating complementary spatial-temporal information across multiple views, our approach generalizes robustly across populations and may enable real-time decision support, extend LVOT assessment to portable or resource-limited settings, and complement Doppler-based evaluation for longitudinal hypertrophic cardiomyopathy management.
O. Crystal, J. Farina, I. Scalia et al.· Circulation Cardiovascular I...· 0 citations
The proposed anatomically guided BackMix augmentation combined with semi-supervised ensemble learning can improve classification accuracy, robustness, and interpretability in echocardiographic analysis under limited annotation conditions, offering a promising approach for automated AS assessment across independent clinical datasets.
Fatima Ezzahra Elkouahy, Badreddine Labakoum, H. Ouahid et al.· Journal of Electronics Elect...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.