Performance of large language models on screening titles and abstracts of urology-related systematic reviews
Abstract
Background: Systematic reviews summarize research evidence to inform clinical practice guidelines and health policy, but the review process is labour-intensive. Authors often screen thousands of abstracts to identify relevant studies to include in their review. Large Language Models (LLMs) may improve title and abstract screening efficiency, but their performance compared to humans require evaluation. Objective: To assess whether locally deployed LLMs accurately screen titles and abstracts of urology-related systematic reviews. Methods: We used LLMs to screen titles and abstracts from three Cochrane reviews in urology totaling 8020 records. Records were screened using predefined criteria, as stated in the Cochrane manuscripts, using three strategies: Qwen3 30B A3B Instruct 2507 alone, Ministral 3 14B Instruct 2512 alone, and their ensemble. A “first-ahead-by-k” approach (k = 2) was used with two models in parallel. We compared their outputs, and once a classification led by a margin of K (i.e., it appeared two more times than any alternative in the running tally) we selected that label as the consensus. Strategy performance metrics were calculated against the human-defined gold standard. Results: Specificity consistently exceeded 90% across strategies (Figure 1). Pooled estimates of sensitivity across the three reviews for strategies were: Ministral 3 14B, 71.1% (95% CI, 64.9-76.6); Qwen3 30B A3B, 53.1% (46.6-59.4); and their ensemble, 72.2% (66.1-77.7). Conclusion: Locally deployed LLMs can screen titles and abstracts with moderate accuracy. Next steps to improve LLM performance include creating random subsamples from each systematic review to perform calibration screening runs with the models, reviewing discrepancies between the model and human-defined gold standard to determine why the performance was variable, and retrieving the exact title and abstract screening criteria used by the authors.