Findings and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection, and local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset.
Abstract
An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report generators. We study Jev, a System One decision model, as a simple, low-cost judge of agreement with physician-written reference reports. Our evaluator checks whether each statement is supported by the other report and combines these judgments in both directions to capture unsupported claims and omissions. A single-question configuration reaches Kendall correlations of 0.573 on RadEvalX and 0.398 on RadEvalExpert with expert error counts, outperforming an open natural language inference judge under matched decomposition and aggregation. One support question per statement retains similar expert agreement to seven while using 43-45% fewer judgment input tokens. At the documented API price, judgments cost under three cents per hundred report pairs, excluding local decomposition. In a separate controlled-error test, Jev detects false negation with an AUROC of 0.977. Local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset. Finding-count and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection. These results support Jev as a practical judgment component for measuring factual differences in generated radiology reports and identify where more elaborate evaluation remains valuable.
RadMatch is the most clinically aligned metric, matching inter-radiologist agreement on ReXVal and more than doubling the best prior metric on the harder RadEvalExpert, and is designed to extend to other modalities and anatomies.
Charles Corbière, Léo Machado, Aubin Charley et al.· 1 citation
Tool-using agents are being proposed for medical imaging, and their behaviour when a tool returns a false finding is largely unmeasured. We audit whether a ReAct-style tool-calling agent abandons an answer it has already given correctly once a falsified finding arrives, and whether that depends on how the finding is pr...
Ridam Roy, Mardin Rashid, Margareta Mia· 0 citations
Medical large language models are evaluated through factuality, physician preference, source support, retrieval quality, safety, and clinical utility. These measures describe content and performance while leaving unresolved whether a generated claim meets the requirements for a specified clinical use. This article prop...
Murat Sariyar· Frontiers in Digital Health· 0 citations
This paper quantifies the sensitivity of established evaluation metrics to variations in reporting practices and introduces a radiologist-informed taxonomy of variations in radiology reporting practice and a method (ReRef) that rewrites reference reports along the axes of the authors' taxonomy while preserving clinical...
Daniel P. Jeong, Charles Q. Li, H. Hosseiny et al.· 0 citations
Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific ques...
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur et al.· 0 citations
It is asked whether judges detect omissions in clinical notes, and two methods reach it independently and trade off: a per-fact pipeline, and a GEPA-evolved prompt doing the same in one call.
Sebastian Fox, L. Markham, Ryan Lail et al.· 0 citations