Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality
Findings and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection, and local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset.