When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation
A taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that creates matched correct and error-injected responses are developed, finding that more specific rubrics better distinguish correct from hallucinated responses.
Griffin Farrow, Lily Sijia Li, Jack Johnson et al.
· 0 citations