Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance
This work systematically investigates how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA), and shows that for supervised image classifiers, changing the label source leads to substantial differences not only in performance estimates but also in model rankings.