Evaluation metrics for synthetic medical imaging.
Abstract
Background
Advances in generative artificial intelligence (AI) have accelerated the development and application of synthetic medical imaging. Despite this rapid progress, the evaluation of synthetic medical images remains heterogeneous, with numerous metrics proposed to assess fidelity, realism, diversity, and clinical validity. Currently, no standardized framework exists to guide the selection, interpretation, or comparison of these metrics, limiting reproducibility and cross-study comparability. This systematic review aims to comprehensively summarize and categorize existing metrics used to assess these complementary dimensions of synthetic medical images.
Methods
A systematic review was conducted in accordance with PRISMA guidelines. PubMed/MEDLINE, EMBASE, Scopus, and arXiv were searched for studies published between 2015 and April 30, 2025, supplemented by citation screening of included studies. Eligible studies were full-text articles that applied or proposed metrics to evaluate the fidelity, realism, diversity, and/or clinical validity in synthetic medical images.
Results
A total of 47 studies were included. Evaluation practices were highly heterogeneous. Expert evaluation (n = 25, 53%) and reference-based evaluations were most common (n = 25, 53%), followed by no-reference metrics (n = 24, 51%), and task-based evaluations (n = 24, 51%). The most commonly used individual metrics were peak signal-to-noise ratio (PSNR) (n = 16, 34%), structural similarity index (SSIM) (n = 15, 32%), mean absolute error (MAE) (n = 12, 26%), and Fréchet Inception Distance (FID) (n = 12, 26%).
Conclusion
Evaluation strategies for synthetic medical imaging showed substantial variability and no single metric captured fidelity, realism, diversity, and clinical validity simultaneously. Metric choice is often dictated by data availability rather than clinical purpose. A task-specific, layered evaluation framework could improve comparability and facilitate clinical adoption.