Jul 2026
Performance of large language models in electrocardiogram interpretation: A comparative study.
The Macro F1 scores of both ChatGPT and Gemini suggest that they are not reliable for independent clinical diagnoses in cardiology, with difficulty shown in interpreting ECGs for rhythm and axis.
Gregory W. Chai, Samuel J. Chen, Jasmine Yang et al.
· Journal of Electrocardiology · 0 citations