Skip to content

Author

Rahmatollah Beheshti

University of Delaware

We have 1 of 64 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Toward Better Assessment of LLMs'Performance in Clinical Error Detection

While models consistently locate error-relevant content, they fail to produce the corresponding correct verdict on the clean counterpart and it is shown that F1 and pairwise accuracy are driven in opposite directions by the same underlying bias, so that ranking models by F1 may systematically promote the weakest discriminators.

Yifan Zhang, Rahmatollah Beheshti · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.