Jul 2026· Journal of Electrocardiology· Vol 98, pp.
154411
· 0 citations· 34 references
Medicine
TL;DR
The Macro F1 scores of both ChatGPT and Gemini suggest that they are not reliable for independent clinical diagnoses in cardiology, with difficulty shown in interpreting ECGs for rhythm and axis.
Abstract
Purpose
As large language models (LLMs) are increasingly used to interpret medical concerns, rigorous evaluation of their performance on clinically relevant tasks is essential. However, the new state-of-the-art models from OpenAI and Google as of February 2026, have not been evaluated for their accuracy and consistency in interpreting ECGs independent of clinical context. We aim to compare ChatGPT (GPT-5.2 Thinking) and Gemini (Gemini 3 Pro) on electrical axis and heart rhythm identification to assess current clinical usability and identify areas for improvement.
METHODOLOGY
ECGs were obtained from the Lobachevsky University Electrocardiography Databases on PhysioNet. The LLM responses were evaluated for first-shot accuracy and consistency across three different trials. First-shot accuracy was further split into the various types of rhythms and axes to examine systematic trends in model outputs.
Results
Both models demonstrated comparable overall first-shot accuracies for electric axis and rhythm classification. Both models struggled with less common rhythm categories, including multifocal rhythms and tachycardias. The Macro F1 analysis indicated low overall classification performance for both ChatGPT and Gemini in terms of axis and rhythm. ChatGPT achieved a higher Macro F1 point estimate for axis classification compared with Gemini, though both models struggled with a Macro F1 of 45.2% and 32.4% respectively.
Conclusion
The Macro F1 scores of both ChatGPT and Gemini suggest that they are not reliable for independent clinical diagnoses in cardiology, with difficulty shown in interpreting ECGs for rhythm and axis. This investigation aims to guide continued improvement of these LLM models for physician assistance.
Automated electrocardiogram (ECG) interpretation has advanced, yet most systems remain narrow classifiers that emit fixed labels rather than the narratives or endpoint-specific answers clinicians need. Generative approaches could instead produce rich narratives, but are constrained by the gap between continuous biosignals and discrete language tokens. Here we present DeepECG-Tok, which reframes ECG interpretation as a unified instruction-following problem. A residual vector-quantization tokenizer (QINCo) maps 12-lead waveforms to language-model-compatible tokens. Its frozen embeddings achieved a macro-averaged area under the receiver operating characteristic curve (AUROC) of 0.96 for 77-condition classification, outperforming supervised and self-supervised baselines. Aligned with a large language model, a single instruction-tuned model performed ECG interpretation, structured reporting and clinical endpoint prediction, including left ventricular ejection fraction, structural heart disease and atrial fibrillation risk, using 7.27 million question answer pairs. Endpoint classifiers using the frozen tokenizer transferred without retraining to four external cohorts, retaining AUROCs of 0.88--0.90. Evaluated end to end using an ontology-grounded large language model as a judge, which achieved a mean agreement (Cohen's K) of 0.82 against two cardiologists, the unified instruction-tuned model scored 0.71 internally and 0.50--0.53 in the same external cohorts. In a blinded reader study, board-certified cardiologists and residents rated its free-text reports comparably to reference clinician reports, with a paired win-tie-loss distribution of 32:35:33 and a forced-choice preference of 0.53 among decided cases, with no significant difference between the model and reference reports in either comparison. These findings establish discrete ECG tokenization as a foundation for general-purpose models that generate clinically useful interpretations and answer diverse questions directly from cardiac waveforms.
R. Banerjee, J. Beaulieu, N. Dostie et al.· medRxiv· 0 citations
Background: Early and accurate identification of ST-Elevation Myocardial Infarction (STEMI) in the prehospital setting would reduce morbidity and mortality (Rao et al. 2025). Artificial Intelligence (AI) tools may assist emergency medical personnel by providing rapid electrocardiogram (ECG) interpretation and differential diagnosis (Chen et al. 2022; Nallamothu et al., 2015).
Objective: To evaluate the diagnostic accuracy of ChatGPT-4.0, Gemini 2.5 Pro, and a hybrid model combining ECG-GPT + ChatGPT-4.0 in classifying STEMI versus non-STEMI cases based on ECG and clinical data inputs.
Methods: Fifty-six consecutive de-identified cases (28 STEMI, 28 non-STEMI) from Staten Island University Hospital EMR (1/2025-6/2025) were analyzed. Each case included demographics, vital signs, chief complaints, and a representative 12-lead ECG. The three AI models were independently asked to classify each case as STEMI or not STEMI.
Results: For the STEMI cohort (n=28), detection rates were: ChatGPT-4.0 21.4%, Gemini 2.5 Pro 67.9%, and ECG-GPT + ChatGPT hybrid 71.4%. For the non-STEMI cohort (n=28), correct classification rates were: ChatGPT-4.0 92.9%, Gemini 2.5 Pro 28.6%, and ECG-GPT + ChatGPT hybrid: 96.4%. The ECG-GPT + ChatGPT hybrid model demonstrated the highest diagnostic accuracy at 83.9%.
Conclusion: While ChatGPT-4.0 showed high specificity, it lacked sensitivity for STEMI detection. Gemini 2.5 Pro improved sensitivity but yielded higher false positives. The hybrid ECG-GPT + ChatGPT model outperformed both standalone models, suggesting that multimodal AI integration may enhance triage accuracy for prehospital care. Despite the small population size and limitations associated with this, these results propose a viable means to hypothesize that AI may have potential in assisting responders with early STEMI identification. Future improvements in model precision could support the implementation of AI-based triage systems in emergency response settings, reducing time to treatment and improving outcomes by aiding direct cardiac catheterization lab transfers without unnecessary delays (Garvey et al., 2012; Squire et al., 2014).
Catherine Royzman, Jonathan Spagnola, Ruben Kandov· International Journal of Par...· 0 citations
While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals. In the critical field of dynamic electrocardiograms (ECG), models struggle with complex temporal reasoning and diagnostic report generation due to a lack of high-quality datasets and benchmarks. To address this, we introduce (i) Holtercare-23K, a large-scale multimodal dynamic ECG dataset comprising 22,980 QA pairs derived from 788 clinical Holter records and featuring a novel signal-video-text tri-modal alignment. Based on this dataset, we present (ii) Holtercare-Bench, a multimodal benchmark that evaluates models on temporal localization, clinical diagnosis, and global summarization. Zero-shot evaluations of leading MLLMs reveal a significant performance gap in processing ultra-long pathological sequences. However, fine-tuning representative models yields substantial improvements. This work illuminates the limitations of current MLLMs in electrophysiology and provides a foundational benchmark for long-term medical MLLMs. Our project is available at https://github.com/ZJU4HealthCare/Holtercare-Bench.
ECGQuest provides a reproducible benchmark for contextual ECG knowledge and shows that parameter-efficient fine-tuning can make smaller language models competitive with substantially larger commercial models.
M. Hassannia, Matthew A. Reyna, R. Sameni· 0 citations
For automated assessment metadata extraction, reliability rather than accuracy determines whether an LLM can be deployed, and national origin is not a meaningful predictor.
Hui Zhang, Lihui Qu, Hongbo Bai et al.· Frontiers in Artificial Inte...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.