Skip to content
Open access

Pan-retinal pathology detection in oct scans integrating natural language synthesis with diagnostic annotation

Aug 2026 · npj Digital Medicine · 0 citations

TL;DR

iOCT is established as a deployable, multi-disease OCT system with performance approaching trained ophthalmologists, supporting large-scale retinal disease screening in primary care and automated multi-sectional scan analysis and comprehensive natural language report generation.

Abstract

While OCT is pivotal for macular disease diagnosis, its adoption in primary care is limited by AI systems that cannot simultaneously analyze multi-sectional scans across the full spectrum of maculopathies or generate diagnostically integrated reports. Here we present iOCT, an intelligent OCT analysis system that integrates a multi-level annotation framework with data distillation to enable automated multi-sectional scan analysis and comprehensive natural language report generation. iOCT was trained on 107,790 macular OCT scans (1,296,439 images) and annotated across four levels combining case-level natural language descriptions with image/study-level diagnostic classifications. Internally, iOCT achieved BLEU-1 of 0.5995 for report generation—outperforming all baselines including R2Gen—and a mean AUC of 0.988 (95% CI: 0.985–0.991) for diagnostic classification. In prospective multicenter validation across ten hospitals (8998 cases), Integrated Reports achieved physician-level quality in 96.4% of cases (mean 2.94/3), significantly outperforming NL Reports (79.5%, 2.60/3), with the greatest gains at previously low-performing centers (1.38–2.93). iOCT matched junior ophthalmologists in speed (23.37 s vs. 24.44 s, p  = 0.368) and report quality (2.91 vs. 2.88, p  = 0.279), outperformed residents ( p  < 0.001), and approached senior specialists (19.47 s, 2.98). These findings establish iOCT as a deployable, multi-disease OCT system with performance approaching trained ophthalmologists, supporting large-scale retinal disease screening in primary care.

Read PDF

Similar papers

Conference Aug 2026

SCAN-R: bridging precision imaging and natural language for automated stroke diagnosis

This work presents Stroke CT Analysis and Natural Language Reporting (SCAN-R), a unified end-to-end framework that integrates multiclass stroke detection, Transformerenhanced U-Net segmentation with task-specific pre-trained backbones, and Retrieval-Augmented Generation for evidence-based clinical report generation.

Le Minh Toan Truong, X. Nguyen, Dang Khanh Tran · 0 citations
Aug 2026

Diagnostic Accuracy of Multimodal Large Language Models in Retinal Fundus Photography.

ChatGPT demonstrated the strongest accuracy, justification-accuracy association, and confidence-accuracy correlation compared with Claude, Gemini, and Grok and showed a positive correlation between confidence and accuracy.

Lia Huo, Astha Chandra, Michael Balas et al. · 0 citations
Open access Sep 2026

Comparative Study On Optical Coherence Tomography (OCT) Retinal Disease Segmentation

A systematic comparative evaluation of three deep learning-based segmentation architectures — U-Net, U-Net++, and Y-Net — for automated identification of DME and Intraretinal Fluid regions in OCT scans demonstrates the feasibility of deep learning-based OCT segmentation as a diagnostic support tool in resource-constrai...

Dhiraj Pyakurel, Yokisha Poudel, Sushiksha Prasai et al. · 0 citations
Review Open access Aug 2026

Vision and Language Models for Classifying Maxillary Sinus Disease on Cone-Beam Computed Tomography: A Transparent Multimodal Benchmark

Background: Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abnormality is clinically relevant. Unlike most vision-language benchmarks in medical imaging, which pair images with pre-existing, human-authored clinical reports, findings...

S. Alhebshi, H. Khalifa, T. D. Pham · 0 citations
Preprint Aug 2026

Automated 2D and 3D Segmentation of AMD and DME Lesions in OCT

Age-related macular degeneration (AMD) and diabetic macular edema (DME) are leading causes of vision loss, and optical coherence tomography (OCT) is the standard modality for detecting and monitoring the subtle lesions that drive treatment decisions. Most deep-learning segmentation work for OCT is validated only in-dom...

Lucia Sundberg, Zhi-Hao Zhao, M. Nasseri · 0 citations

Can large language models unlock discrete data in ophthalmic diagnostic reports?

Objective: To assess the accuracy and efficiency of a large language model (LLM) using two prompt strategies to extract structured data from ophthalmic diagnostic PDF reports. Methods: Twenty deidentified reports across four types (Visual Field, OCT Glaucoma Overview, OCT retinal nerve fiber layer Single Exam, and OCT...

Umair A. Zaidi, An-Lun Wu, Wei-Chun Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.