Skip to content

Author

Abraham Tsur

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Multimodal large language models for bladder tumor detection in cystoscopy: a retrospective benchmarking study

Cystoscopic assessment is central to bladder cancer diagnosis, yet visual interpretation remains variable. Existing artificial intelligence approaches often depend on data-intensive models that are often difficult to deploy in routine practice. We evaluated whether multimodal large language models (MLLMs), including smaller and more efficient architectures, can accurately classify cystoscopy images, and whether prompt engineering improves performance. The primary outcome was benign-versus-malignant classification. Secondary outcomes included calibration, high-confidence triage, and performance stratified by imaging modality. We retrospectively analyzed 1,754 labeled public cystoscopy images. Three prompt types were tested: Direct, Book-based, and Optimized, across GPT-5.2, GPT-5, GPT-5-Mini, and GPT-5-Nano. Performance measured: accuracy, sensitivity, specificity, and F1 Score. Confidence evaluation: using Brier Score and Expected Calibration Error. High-confidence triage using abstention option based on loss function. GPT-5 and GPT-5-Mini with the optimized prompt achieved the best benign-versus-malignant performance, with accuracies of 86.7% and 89.2%, specificities of 94.1% and 88.4%, and sensitivities of 82.5% and 89.2%, respectively. GPT-5 with the optimized prompt achieved the best high-confidence triage performance, yielding 98.1% accuracy, 94.6% specificity, and 99.1% sensitivity at 62.0% image coverage. Prompt engineering improved model performance, although these gains were not statistically significant, and enhanced confidence calibration and triage performance. This retrospective evaluation demonstrates the potential of MLLMs for cystoscopic bladder lesion classification. Prompt engineering improved diagnostic calibration and output reliability, while high-confidence triage increased accuracy to 98.1%, supporting the feasibility of MLLMs as foundation models for cystoscopic assessment.

Yonatan Prat, Husny Mahmud, Abraham Tsur et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.