Deep learning models can support detection and classification of hiatal hernias using focused AP barium swallow images, even when trained on a limited dataset, and this training approach may enable automated identification and classification while potentially reducing radiation exposure and barium ingestion during diagnostic studies.
Abstract
Benign Disease: New Technologies
Hiatal hernia diagnosis using barium swallow studies often requires multiple image acquisitions to visualize the esophagogastric junction adequately. Repeated acquisitions may increase radiation exposure and barium ingestion. Interpretation is observer-dependent. The aim of this study was to train an artificial intelligence model for detecting and classifying hiatal hernias.
A limited dataset of 70 anonymized barium swallow images centered on the esophagogastric junction was retrospectively analyzed and classified as hiatal hernia (n=45) or normal (n=25). Images were standardized using region-of-interest cropping, autocontrast adjustment, resizing to 512×512 pixels, and data augmentation to mitigate small-sample limitations. The dataset was divided into training (70%), validation (15%), and testing (15%) subsets.
Three pretrained convolutional neural networks (ResNet18, DenseNet121, EfficientNetB0) were fine-tuned using transfer learning. A structured grid search (72 trials) optimized learning rate, dropout, weight decay, and training epochs under identical cross-validation conditions. The primary evaluation metric was mean area under the curve (AUC). Secondary metrics included F1-score, accuracy, sensitivity, and specificity. Grad-CAM visualization was applied to assess anatomical regions influencing predictions.
Despite the limited dataset and use of focused images, all architectures demonstrated strong internal performance. DenseNet121 achieved mean AUC and F1 values of 1.00 in cross-validation, while EfficientNetB0 showed the strongest out-of-fold performance (AUC 0.9733; F1 0.9556). ResNet18 achieved mean cross-validation AUC 0.9956 and F1 0.9882.
Under a unified grid-search comparison framework, ResNet18 demonstrated the most consistent global performance (mean AUC 0.7860; F1 0.7662; accuracy 0.7571; sensitivity 0.6215; specificity 1.0000), supporting its selection as the final model. Grad-CAM confirmed consistent attention to the esophagogastric junction region. The prototype web application generated automated binary classification with confidence scoring.
Deep learning models can support detection and classification of hiatal hernias using focused AP barium swallow images, even when trained on a limited dataset. This training approach may enable automated identification and classification while potentially reducing radiation exposure and barium ingestion during diagnostic studies.
A Vision Transformer deep learning system is developed and prospectively validated to predict reduction failure from static B-mode ultrasound images to assess pediatric ileocolic intussusception severity, providing an objective, accurate tool to support clinical decision-making and reduce treatment risks.
Jie Liu, Yue Wang, Danping Zeng et al.· npj Digital Medicine· 0 citations
The proposed framework establishes a reliable and lightweight baseline for automated gastrointestinal disease detection and demonstrates that ConvNeXt-Tiny effectively captures disease-relevant visual patterns in endoscopic images while maintaining consistent performance across varying training conditions.
Muhammad Faqih, O. Q. Aziz, Ajib Hanani· Jurnal Ilmu Komputer dan Inf...· 0 citations
An advanced system capable of automatically detecting complicated appendicitis from ultrasound images was developed and was explained with gradient-weighted class activation mapping (Grad-CAM), which creates a heatmap of the regions responsible for the model's prediction of the infected areas.
Fahad Ahammed, Omar Faruq Shikdar, Navid Zaman et al.· 0 citations
Random-split evaluation substantially overestimated performance relative to a resolution-defined acquisition-shift stress test, and entropy-based selective prediction improved reliability by identifying a high-confidence subset for automated prediction while deferring the remainder to human review.
Behnam Kiani Kalejahi, S. Khan, Murodbek Akhrorov et al.· Biomedicines· 0 citations
While the results are promising for a novel application domain, the model's failure on clinically critical minority classes (Bilateral Blockage, Bilateral Patency) means it is not yet suitable for unsupervised clinical use.
Nasreen Jawaid, I. Brohi, Najma Imtiaz Ali et al.· Scientific Reports· 0 citations
The feasibility of developing a deep learning model for the detection of intussusception using a smaller dataset of POCUS images is demonstrated and fine-tuning models were best adapted to the screening nature of POCUS images.
A. Thyagachandran, Brian Lefchak, H. Murthy et al.· Frontiers in Radiology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.