The Ability of Generative Artificial Intelligence to Evaluate the Quality of Scholarly Output
Large language models (LLMs) are increasingly being considered as tools to support the peer-review process. However, we currently have a limited understanding of their capabilities and biases in this area. In a set of three studies, we examined how out-of-the box LLMs evaluate scientific quality when presented with con...