Skip to content
Open access

Accuracy Evaluation of LLM-Generated Electronic Health Record Interpretations Against Medically Verified Sources

Sep 2026 · Bioengineering · 0 citations · 23 references

Abstract

This study evaluates the accuracy of medical-report interpretations generated by large language models in comparison to medically verified sources, with a particular focus on urology. The main goal is to examine the extent to which LLMs can reliably and precisely explain medical findings, with emphasis on expert urological terminology. For the evaluation of interpretation accuracy, real clinical documentation was used, specifically a dataset of 40 urological reports from the Clinical Hospital Center Rijeka. The generated explanations were compared with medically correct definitions using quantitative metrics assessing interpretability, specificity, and clinical accuracy. The results revealed a clear stratification of model performance, with DeepSeek-V3.2, GLM-4.6, MiMo-V2-Flash, and Qwen2.5-7B-Instruct achieving the highest semantic accuracy and consistency, while several models exhibited unstable behavior characterized by occasional catastrophic failures. Overall, the findings indicate that although LLMs can produce clinically coherent explanations of urological findings, their variability and susceptibility to hallucinations necessitate human oversight, supporting their use primarily as decision-support tools rather than autonomous clinical interpreters.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.