Skip to content

Author

Holger Kunz

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Automated chest X-ray disease screening using large language models and deep convolutional neural networks on the MIMIC-CXR dataset

Radiologists in resource-limited settings often face high workloads, especially in chest X-ray interpretation. Manual annotation of large-scale imaging datasets remains costly and time-consuming. This study aims to explore the feasibility of using large language models (LLMs), specifically GPT-4o, to generate binary disease presence labels from free-text radiology reports, and to use these labels to train deep learning models for automated chest X-ray classification. A two-stage supervised learning pipeline was developed using the publicly available MIMIC-CXR v2.1.0 dataset. First, GPT-4o was prompted with a structured clinical protocol to classify each radiology report as either “diseased” or “no disease.” Second, the generated labels were used to supervise the training of four convolutional neural networks: ResNet-18, DenseNet-121, EfficientNet-B1, and ConvNeXt-Tiny. A patient-level 70/10/20 split was employed to prevent data leakage across sets. Each model was trained across five random seeds (42–46), and 95% confidence intervals were computed using the t-distribution. Label quality was evaluated by comparing 210 generated labels against radiologist annotations from a board-certified radiologist. GPT-4o achieved an overall accuracy of 92.9% with expert labels on the 210-report validation set. For the “diseased” class, the precision was 97.4% and recall was 90.5%; for “no disease,” precision was 87.1% and recall was 96.4%. Among the CNN models evaluated on the held-out test set, ConvNeXt-Tiny achieved the highest area under the curve ( AUC = 0.832 , 95% CI [0.801, 0.863]) and balanced accuracy (0.739), significantly outperforming EfficientNet-B1 ( AUC = 0.797 ; paired t -test, p = 0.014 ). ResNet-18 ( AUC = 0.822 ) and DenseNet-121 ( AUC = 0.808 ) showed intermediate performance. All models demonstrated AUC values above 0.79, confirming the viability of LLM-derived weak supervision. This study demonstrates that LLMs can be effectively employed to generate supervision labels for medical imaging tasks. The proposed approach offers a scalable and low-cost solution for preliminary disease screening, particularly in healthcare environments with limited expert availability. The multi-seed evaluation with confidence intervals provides a rigorous assessment of model stability. Further work is needed to improve label reliability and expand to multi-label classification.

Qing-Yuan Zhang, Pardeep Vasudev, Kezhi Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.