Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages
Across this controlled multilingual corpus, aggregate transcription-frequency metrics did not reliably track the sparse severe tail of clinically consequential errors, and these data do not support its use alone as a proxy for clinical safety.