This paper quantifies the sensitivity of established evaluation metrics to variations in reporting practices and introduces a radiologist-informed taxonomy of variations in radiology reporting practice and a method (ReRef) that rewrites reference reports along the axes of the authors' taxonomy while preserving clinical interpretation.
Abstract
Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in terminology, shorthand, formatting, and level of detail. These variations in reporting norms represent an under-appreciated obstacle in efforts to evaluate AI-based radiology report generation (RRG) models, where machine-generated reports are typically assessed based on their concordance with human-generated references. In this paper, we quantify the sensitivity of established evaluation metrics to variations in reporting practices, revealing impacts large enough to alter the rankings of models. We introduce a radiologist-informed taxonomy of variations in radiology reporting practice and a method (ReRef) that rewrites reference reports along the axes of our taxonomy while preserving clinical interpretation. For instance, when comparing the performance of nine RRG models on MIMIC-CXR using RadCliQ-v1, condensing the discussion of normal findings in the reference reports causes Libra to drop from first to second place while CheXOne rises from third to first. Our results suggest that many current metrics fail to decouple clinical interpretation from conformity to reporting practices and that choosing the ``right''references that accurately reflect the desired reporting practices can be important in practice. To support future research, we release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 (original, alternative) reference report pairs derived from MIMIC-CXR.
We developed a national questionnaire to assess radiology reports (RR) physicians’ reading habits and identify their expectations towards RR, focusing on CT and MRI reports. An anonymized 20-item questionnaire was nationally distributed in France via practice and hospital mailing lists and one social media group, and i...
Yasmine Kassab, Matthieu Bailly, Sixtine Brabant et al.· Insights into Imaging· 0 citations
RadMatch is the most clinically aligned metric, matching inter-radiologist agreement on ReXVal and more than doubling the best prior metric on the harder RadEvalExpert, and is designed to extend to other modalities and anatomies.
Charles Corbière, Léo Machado, Aubin Charley et al.· 1 citation
Errors in radiology reports are a major patient-safety concern and are difficult to detect with manual quality assurance (QA). Large language models (LLMs) can assist, but generic prompting does not reflect radiologists’ structured, section-based workflows. To develop and evaluate RadCoT (Radiological Chain-of-Thought)...
Jia Li, Zi-Chun Zhou, Yan-Tao Niu et al.· European Radiology Experimen...· 0 citations
This Educational Editorial & Practical Guide translates CARE-radiology, CARE, ICMJE, COPE, the 2024 Declaration of Helsinki and CRediT into a 12-step SJORANM workflow for authors, reviewers and editors to create an educational, internally consistent and independently verifiable case report before peer review.
F. Mosler, G. Nöldge, Keivan Daneshvar· Swiss Journal of Radiology a...· 0 citations
Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. Such control is essential because clinical scenarios diverge: emergen...
Ying Jin, Noel C. F. Codella, John Corring et al.· 0 citations
This study evaluates the accuracy of medical-report interpretations generated by large language models in comparison to medically verified sources, with a particular focus on urology. The main goal is to examine the extent to which LLMs can reliably and precisely explain medical findings, with emphasis on expert urolog...
Mia Rovis, Alma Smajić, Marijela Miličević et al.· Bioengineering· 0 citations