Skip to content
Open access

Comprehensive Echocardiography Interpretation Using Video and Multiview Vision-Language AI.

Jul 2026 · JACC: Asia · 0 citations · 17 references
Medicine

TL;DR

A multiview video-language framework improved report retrieval compared with conventional image-based approaches and support the utility of video-based, multiview representation learning for echocardiographic report retrieval.

Abstract

Background

Echocardiography is essential for assessing cardiac structure and function, yet accurate interpretation requires specialized expertise, creating challenges in emergency care and in regions with limited access to experienced echocardiographers. Although artificial intelligence-based interpretation is increasingly studied, most existing models rely on still images or single views and do not reflect the video-based, multiview integration used in clinical practice.

Objectives

The authors aim to develop and evaluate a multiview video-language model that integrates cardiac motion and multiple standard echocardiographic views.

Methods

We trained a vision-language model on 577,061 transthoracic echocardiography videos paired with Japanese clinical reports from 46,852 examinations acquired at a single tertiary center (2015-2023). For each examination, features from 5 standard views (parasternal long-axis and short-axis; apical 2-, 3-, and 4-chamber) were aggregated. Performance was assessed in a retrieval task selecting the correct report from 16,436 candidate reports in the test cohort based on video input; the primary metric was correct match retrieval probability within the top 10 candidates (ie, R@10).

Results

The image-based baseline achieved an R@10 of 1.3%. Using video input increased R@10 to 4.2%. Integrating 5 views further improved R@10 to 6.0% (95% CI: 5.5%-6.4%). Gains were greatest for findings that depended on temporal dynamics and multiview assessment, including left ventricular wall motion abnormality and left ventricular dilation.

Conclusions

A multiview video-language framework improved report retrieval compared with conventional image-based approaches. These findings support the utility of video-based, multiview representation learning for echocardiographic report retrieval.

Read PDF

Similar papers

Open access 2026

ROI-Guided Echocardiographic Image Analysis and Case Aggregation for Differentiating Atrial and Ventricular Septal Defects

This study implemented strict subject-disjoint partitioning to eliminate data leakage, and simultaneously introduced cross-frame case aggregation to emulate the multi-frame visual synthesis process of expert echocardiographers, suggesting that the proposed workflow has the potential to serve as an adjunctive tool for septal defect screening.

Tao Zhang, Peipei Zhang, Qing-Yuan Zhang et al. · 0 citations
Aug 2026

Beyond Doppler: Scalable AI Detection of LVOT Obstruction in HCM.

BACKGROUND Accurate assessment of left ventricular outflow tract (LVOT) gradients is critical for hypertrophic cardiomyopathy management, yet Doppler-based measurements are technically demanding and require expertise. The objective of this work was to develop a multi-view deep learning model capable of classifying LVOT obstruction (>20 mm Hg) using routine 2-dimensional echocardiographic windows without reliance on Doppler imaging. METHODS We trained and externally validated a cross-attention-based video-to-video fusion framework that integrated EchoPrime-derived video representations from 3 standard transthoracic echocardiographic views to classify LVOT gradients. RESULTS Training was performed on a derivation cohort (N=1833) from a tertiary care system in the United States, with model performance evaluated on an internally held-out test set (N=275) and a Korean external validation cohort (N=46). Single-view baselines showed limited discrimination (external area under the receiver operating curves, 0.47-0.70). Conversely, the domain-specific foundational model (EchoPrime) achieved superior single-view performance (area under the receiver operating curves, 0.75-0.80 internal; 0.79-0.83 external), highlighting the importance of echo-specific pretraining and temporal modeling. The proposed multi-view fusion further enhanced predictive performance, with the late fusion model reaching an area under the receiver operating curve of 0.84 on the external cohort with significant population-shift. CONCLUSIONS These results suggest LVOT physiology is encoded in routine 2-dimensional imaging and can be leveraged for clinically relevant gradient classification without Doppler input. The proposed artificial intelligence-guided strategy demonstrates substantial cost savings compared with the screen-all approach. By integrating complementary spatial-temporal information across multiple views, our approach generalizes robustly across populations and may enable real-time decision support, extend LVOT assessment to portable or resource-limited settings, and complement Doppler-based evaluation for longitudinal hypertrophic cardiomyopathy management.

O. Crystal, J. Farina, I. Scalia et al. · 0 citations
Open access Sep 2026

Point-of-care echocardiography screening for hypertrophic cardiomyopathy using automated deep-learning analysis

Abstract Aims Hypertrophic cardiomyopathy (HCM) remains underdiagnosed due to limited access to expert imaging. We developed and validated a deep-learning (DL)-based echocardiographic model adaptable to point-of-care ultrasound (POCUS) for scalable HCM screening. Methods and results We retrospectively analysed 134 956 expert transthoracic echocardiograms (TTE) from 73 598 patients at Sheba Medical Center (2007–2022). A TTE-trained DL model integrating structural features and temporal motion patterns from parasternal long-axis and apical four-chamber views estimated HCM probability. Performance was evaluated in an independent test cohort and clinical subgroups. External validation used bedside POCUS studies from non-cardiologists with handheld devices. The test cohort included 12 096 patients with 119 confirmed HCM cases (prevalence 0.98%; median age 75 years, 57% male). HCM-positive patients showed increased expert TTE-measured septal (1.67 [1.5, 2.0] vs. 1.01 [0.9, 1.19] cm) and posterior wall thickness (1.1 [1.0, 1.3] vs. 0.9 [0.8, 1.0] cm) (P < 0.001). The model achieved excellent discrimination with an area under the curve of 0.982 (95% CI 0.966–0.993), sensitivity 88.2%, and specificity 97.3%, robust across subgroups. The POCUS cohort (n = 1047, median age 73 years, 55% male) represented multimorbid inpatients with 65 (6.2%) classified as screen-positive by the algorithm. These showed higher expert TTE-measured septal thickness (1.26 [1.07, 1.46] vs. 1.06 [0.9, 1.2] cm; 22% vs. 4% with IVS ≥1.5 cm; P ≤ 0.01). Among 49 (75%) POCUS-flagged positive patients with formal TTE and clinical data, 8 (16%) were confirmed by expert adjudication to have HCM. Specificity is limited by occasional confounding amyloidosis detection (4% of POCUS-flagged patients). Conclusion This DL-based model identifies HCM and demonstrates feasibility for POCUS screening, supporting earlier detection and broader diagnostic access.

N. Karra, Y. Klempfner, Viana Copeland et al. · 0 citations
Review Open access Sep 2026

Artificial Intelligence in Echocardiography and Point-of-Care Ultrasound: Applications, Clinical Integration, and Future Directions

Artificial intelligence (AI) is increasingly used in echocardiography and point-of-care ultrasound (POCUS) to support image acquisition, view recognition, image-quality assessment, segmentation, automated quantification, disease classification, reporting, and bedside decision support. This narrative review summarizes clinically relevant applications, with emphasis on clinical integration, pediatric and congenital heart disease considerations, and safe implementation. The strongest clinical evidence supports automated left ventricular segmentation and ejection fraction estimation, AI-guided acquisition, and workflow efficiency. Video-based deep learning has enabled beat-to-beat assessment of ventricular function, and a randomized workflow trial showed that AI-generated initial ejection fraction assessment was noninferior to sonographer assessment and required fewer cardiologist corrections. Regulatory-authorized acquisition and analysis tools demonstrate growing clinical adoption for specified adult indications. Recent multiview and disease-phenotyping models extend AI toward more comprehensive interpretation, while AI-enabled POCUS may improve focused image acquisition by non-expert users. However, external validation, pediatric and congenital heart disease data, cross-device generalizability, clinical outcome evidence, uncertainty communication, automation bias, and medicolegal responsibility remain important limitations. AI should currently be viewed as an augmentative rather than autonomous technology. The safest near-term model is human-AI collaboration, in which validated tools improve acquisition, reproducibility, and workflow while clinicians retain responsibility for interpretation and patient-centered decisions. Pediatric and congenital heart disease applications require age- and anatomy-specific datasets, multicenter validation, local performance monitoring, and clinician-supervised deployment. AI can assist across the echocardiography workflow, from acquisition guidance and view recognition to segmentation, quantification, disease screening, and structured reporting. The most mature clinical evidence supports left ventricular segmentation, ejection fraction estimation, AI-guided acquisition, and workflow efficiency rather than autonomous diagnosis. AI-enabled POCUS may improve acquisition by non-expert users, but image adequacy, interpretation, and clinical integration must remain clinician-supervised. Pediatric and congenital heart disease applications are promising but remain less mature than adult applications and require anatomy-specific datasets and multicenter validation. Safe implementation requires external validation, local monitoring, bias assessment, uncertainty display, audit trails, and institutional governance.

F. Savorgnan, Pranathi Pilla, Sarah Visokay et al. · 0 citations
Review Open access Jul 2026

The Future of Imaging in Heart Failure: Toward Precision Phenotyping, Integration, and Intelligence

Heart failure (HF) is increasingly understood not as a single, uniformly treated diagnosis but as a heterogeneous syndrome requiring aetiological clarification, in which cardiac imaging is central. As the opening article of this journal's ‘Imaging in Heart Failure’ section, this review surveys the technologies currently reshaping HF imaging and sets out the section's scope and priorities, framing the shift from a descriptive, modality-siloed practice toward an integrated, predictive, patient-specific discipline. Artificial intelligence (AI) now delivers expert-level echocardiography automation, guides image acquisition by novices in resource-limited settings, detects aetiologies such as transthyretin amyloid cardiomyopathy from a single acquisition and enables deep phenotyping through radiomics and vendor-agnostic strain analysis. Handheld, AI-enabled point-of-care ultrasound extends imaging-guided triage beyond the echocardiography laboratory. Cardiovascular magnetic resonance (CMR) advances — parametric mapping, four-dimensional flow, diffusion tensor imaging, spectroscopy, and accelerated reconstruction — broaden tissue and metabolic characterisation, including patients with implanted devices. Molecular imaging with novel positron emission tomography tracers and hyperpolarised magnetic resonance is moving from depicting the structural consequences of disease to imaging active pathobiology, while photon-counting computed tomography and image-derived digital twins support one-stop structural assessment and in-silico prediction of therapy response. The convergence of AI, molecular imaging and advanced precision is transforming HF imaging from better pictures into smarter, integrated, personalised data that directly inform care. Realising this promise will require rigorous validation, attention to algorithmic bias and generalisability, demonstrated cost-effectiveness, curricular reform, and equitable access. This section aims to critically appraise these innovations and their translation into practice.

M. Hundertmark · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.