Skip to content

Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

Sep 2026 · 6 citations · ⚡ 1 influential · 22 references
Computer Science

TL;DR

Findings and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection, and local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset.

Abstract

An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report generators. We study Jev, a System One decision model, as a simple, low-cost judge of agreement with physician-written reference reports. Our evaluator checks whether each statement is supported by the other report and combines these judgments in both directions to capture unsupported claims and omissions. A single-question configuration reaches Kendall correlations of 0.573 on RadEvalX and 0.398 on RadEvalExpert with expert error counts, outperforming an open natural language inference judge under matched decomposition and aggregation. One support question per statement retains similar expert agreement to seven while using 43-45% fewer judgment input tokens. At the documented API price, judgments cost under three cents per hundred report pairs, excluding local decomposition. In a separate controlled-error test, Jev detects false negation with an AUROC of 0.977. Local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset. Finding-count and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection. These results support Jev as a practical judgment component for measuring factual differences in generated radiology reports and identify where more elaborate evaluation remains valuable.

View source

Similar papers

Preprint Sep 2026

RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching

RadMatch is the most clinically aligned metric, matching inter-radiologist agreement on ReXVal and more than doubling the best prior metric on the harder RadEvalExpert, and is designed to extend to other modalities and anatomies.

Charles Corbière, Léo Machado, Aubin Charley et al. · 1 citation
Preprint Aug 2026

Source-Dependent Deference in Medical Imaging Agents Under Falsified Findings: A Pilot Audit

Tool-using agents are being proposed for medical imaging, and their behaviour when a tool returns a false finding is largely unmeasured. We audit whether a ReAct-style tool-calling agent abandons an answer it has already given correctly once a falsified finding arrives, and whether that depends on how the finding is pr...

Ridam Roy, Mardin Rashid, Margareta Mia · 0 citations
Review Open access Sep 2026

Clinical warrant in medical LLM evaluation

Medical large language models are evaluated through factuality, physician preference, source support, retrieval quality, safety, and clinical utility. These measures describe content and performance while leaving unresolved whether a generated claim meets the requirements for a specified clinical use. This article prop...

Murat Sariyar · 0 citations
#artificial intelligence Preprint Sep 2026

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

This paper quantifies the sensitivity of established evaluation metrics to variations in reporting practices and introduces a radiologist-informed taxonomy of variations in radiology reporting practice and a method (ReRef) that rewrites reference reports along the axes of the authors' taxonomy while preserving clinical...

Daniel P. Jeong, Charles Q. Li, H. Hosseiny et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Scaling Clinical Judgment to Evaluate Medical AI

Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific ques...

Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.