Aug 2026· Frontiers in Neurology· Vol 17· 0 citations· 90 references
Medicine
TL;DR
DL has demonstrated strong potential to improve clinical decision-making and healthcare efficiency in otology, however, the field faces challenges related to data scarcity for rare diseases, poor segmentation performance for tiny structures, and a lack of integration with clinical workflows.
Abstract
Objective Deep learning (DL), a core branch of artificial intelligence (AI), has revolutionized medical image analysis. Driven by advances in computational power and access to large-scale datasets, DL excels at extracting hierarchical features from complex, unstructured data. Accurate interpretation of imaging features is essential for diagnosing and treating middle and inner ear diseases. This narrative review aims to summarize current literature on the application of DL in otological imaging. Methods A narrative literature search was conducted using the PubMed and MEDLINE databases. Keywords related to AI or DL in otology were used to identify relevant articles published up to the time of writing. Results DL models have achieved favorable results in classifying common middle ear diseases (e.g., otitis media) using otoscopic images and in segmenting major inner ear structures (e.g., cochlea, ossicular chain) in 3D volumetric data. While 2D CNNs are mature for otoscopic classification, 3D U-Net and UNETR architectures dominate CT and MRI analysis. Models also show value in low-dose CT reconstruction and multimodal diagnosis. Conclusion DL has demonstrated strong potential to improve clinical decision-making and healthcare efficiency in otology. However, the field faces challenges related to data scarcity for rare diseases, poor segmentation performance for tiny structures (e.g., stapes), and a lack of integration with clinical workflows. Future efforts should focus on standardizing data and optimizing network structures for specific modalities.
Early diagnosis of ear diseases is essential to prevent serious complications such as chronic infections and permanent hearing loss. Conventional otoscopic diagnosis relies heavily on experienced clinicians and specialized equipment, which limits healthcare accessibility in rural and underserved regions. This paper presents an automated ear disease detection system based on Convolutional Neural Networks (CNNs) for binary classification of otoscopic images into normal and abnormal categories. The proposed framework incorporates image preprocessing and data augmentation techniques to improve image quality, increase dataset diversity, and enhance the robustness of the classification model. The model was trained and evaluated using a publicly available dataset containing 1,370 otoscopic images. Experimental results demonstrate an overall classification accuracy of 91.2%, with a precision of 90.4%, a recall of 89.8%, and an F1-score of 90.1%. The proposed system effectively distinguishes normal ear conditions from abnormalities, including Acute Otitis Media (AOM) and other infectious ear diseases. Furthermore, the trained CNN model was integrated into a Flask-based web application to provide an accessible, real-time diagnostic support tool for primary healthcare settings. The proposed solution has the potential to assist clinicians in early screening, improve diagnostic efficiency, and enhance healthcare accessibility in resource-limited environments.
Rakshitha N Poojary, Raksha V Shetty· 2026 International Conferenc...· 0 citations
BACKGROUND
Skin cancer is one of the most common malignancies worldwide, and early detection is essential for improving treatment outcomes and reducing mortality. Conventional diagnostic approaches rely on visual examination and dermoscopic analysis by dermatologists, which can be time-consuming and subject to inter-observer variability. Recent advances in artificial intelligence have enabled the development of computer-aided diagnostic systems to support clinical decision-making in dermatological oncology.
OBJECTIVE
In this study, a deep learning-based framework is proposed for multi-class classification of cutaneous lesions using dermoscopic images.
METHODS
The proposed model integrates a Vision Transformer (ViT) to capture global contextual features and a Squeeze-and-Excitation Residual Network (SE-ResNet) to extract channel-refined local features. These complementary representations are combined using an adaptive attention-based fusion mechanism to improve classification performance. In addition, an Improved Crocodile Optimization Algorithm (ICOA) is employed to optimize model hyperparameters and enhance convergence stability.
RESULTS
The proposed method was evaluated using the HAM10000, ISIC 2019, and PH2 datasets, achieving classification accuracies of 99.05%, 98.31%, and 99.17%, respectively.
CONCLUSION
The results demonstrate the robustness and generalizability of the proposed framework, highlighting its potential as a clinical decision-support tool for early detection and improved diagnosis of skin cancer.
S. Ravisankar, M. Karthikeyan· Cutaneous and Ocular Toxicol...· 0 citations
Artificial intelligence (AI) has emerged as a promising tool in endodontic diagnostics, particularly in radiographic interpretation and treatment outcome prediction. This narrative review provides a comprehensive overview of current applications of machine learning (ML) and deep learning (DL) algorithms in endodontics, focusing on periapical lesion detection, assessment of root canal morphology, prognosis prediction, and navigation-guided procedures. A literature search was conducted in PubMed, Scopus, and Web of Science for studies published between 2021 and 2026. Following qualitative screening, 45 studies were included in the evidence synthesis. Given the narrative design of the review, no formal meta-analysis was performed. The reviewed evidence indicates that convolutional neural networks (CNNs), U-Net architectures, YOLO-based systems, and other AI models have demonstrated promising diagnostic performance across multiple imaging modalities, with the highest performance generally reported for CBCT-based applications. However, the current body of evidence remains limited by retrospective study designs, small datasets, methodological heterogeneity, and insufficient external validation. Overall, AI should be considered a valuable adjunctive decision-support tool in endodontic diagnostics rather than a replacement for clinical expertise. Further prospective multicenter studies, external validation, and standardized methodological frameworks are required to support its reliable integration into routine clinical practice.
M. Radwański, E. Zmysłowska-Polakowska, T. F. Eyüboğlu et al.· Applied Sciences· 0 citations
Skin cancer is one of the most prevalent and deadly malignancies, necessitating early and precise diagnosis to improve patient survival rates. While recent advancements in deep learning have produced robust computer-aided diagnosis (CAD) systems, the majority of these models rely exclusively on visual features extracted from dermoscopic or clinical images. Conversely, dermatologists synthesize visual cues with patient clinical metadata (e.g., age, gender, continuous bleeding, and itchiness) to reach an accurate diagnosis. Prior research attempting multimodal fusion has largely depended on black-box concatenation-based late-fusion strategies or simple neural compression modules, obscuring clinical reasoning. To bridge this gap toward Explainable Artificial Intelligence (XAI), we propose CrossMeta-ViT, a deep vision-transformer metadata-fusion framework designed for transparent skin lesion classification. The core contribution of this architecture is a Cross-Attention Fusion module. Instead of sequentially concatenating features, CrossMeta-ViT utilizes encoded patient tabular metadata as a Query (Q) to dynamically attend to visual image patches acting as Keys (K) and Values (V). This mechanism ensures that structural image features are selected and weighted under the direct guidance of the patient's clinical history. We evaluated the framework on the smartphone-captured PAD-UFES-20 dataset for binary classification (Benign vs. Malignant). Using 5-fold cross-validation, CrossMeta-ViT achieved a macro F1-score of 0.8723 ± 0.0009, malignant recall of 0.9927, specificity of 0.7651, and an AUC of 0.955, while the held-out single-split evaluation yielded an accuracy of 0.90. These results support the model as a recall-oriented and interpretable multimodal aid for telemedicine pre-triage rather than as a universally dominant classifier.
Muhammetalp Erdem· Alfa Mühendislik ve Uygulama...· 0 citations
Ocular pathologies are a leading cause of visual impairment in cats and require accurate, timely diagnosis for effective treatment. This study aimed to comparatively evaluate the performance of different deep learning (DL) architectures for the automated identification and classification of feline ophthalmic disease. A dataset of 456 clinically validated ocular images, spanning 13 classes (12 ocular diseases and a healthy group), was used to train nine convolutional neural network (CNN) architectures, including ResNet, EfficientNet, DenseNet, MobileNet, and VGG models, through transfer learning with pre-trained ImageNet weights. Data augmentation and five-fold cross-validation were applied during training. Statistical significance of performance differences across architectures was assessed using Cohen's Kappa and McNemar's test. On the held-out test set, EfficientNet-B0 reached the highest classification accuracy (78%), while DenseNet-121 obtained the highest macro-averaged F1-score (0.76) and Area Under the Curve (AUC) (0.9790), followed by EfficientNet-B0 (0.9750) and ResNet-34 (0.9724). Confusion matrix analysis showed the strongest classification consistency for cherry eye, healthy, and corneal sequestration (precision and recall reaching 1.00 in several cases), and the weakest for glaucoma (recall 0.25) and corneal ulcer (F-1 score 0.29). McNemar's test indicated no statistically significant difference among EfficientNet-B0, DenseNet-121, and ResNet-34 (p>0.05), and per-image inference times across all nine architectures ranged from 2.67 to 5.73 ms. Gradient-weighted Class Activation Mapping (Grad-CAM) visualizations applied to the EfficientNet-B0 model confirmed that the model focused on clinically relevant ocular regions during classification. Overall, the results obtained with EfficientNet-B0, DenseNet-121, and ResNet-34 suggest that DL-based multi-class classification could support clinical decision-making in the diagnosis of feline ophthalmic diseases.
Vildan Aslan Canatan, U. Canatan, Gizem Kıymet Sancaktar et al.· The Veterinary Journal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.