Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects
This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models'training cutoffs.
Emad Alharbi
· 0 citations