Toward Scalable Audio Description Quality Control: A Workflow for Evaluating Human and VLM Raters
This work developed a methodological workflow using Item Response Theory to evaluate VLM and human rater proficiency against expert-established ground truth, suggesting that top-performing VLMs can approximate ground-truth ratings at levels comparable to human raters.
Lana Do, Gio Jung, J. F. Barajas et al.
· 0 citations