Toward Better Assessment of LLMs'Performance in Clinical Error Detection
While models consistently locate error-relevant content, they fail to produce the corresponding correct verdict on the clean counterpart and it is shown that F1 and pairwise accuracy are driven in opposite directions by the same underlying bias, so that ranking models by F1 may systematically promote the weakest discriminators.