Skip to content

Author

S. Yildirim

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Jul 2026

A Data Curation Framework for Unstructured Real-World Turkish Breast Imaging Reports

Breast imaging reports are typically stored as unstructured free-text documents, which limits their use in clinical analytics, research, and downstream computational applications. These challenges are particularly pronounced in Turkish because of its agglutinative linguistic structure, orthographic variability, and the limited availability of standardized clinical corpora. This study presents a data curation framework for transforming heterogeneous real-world Turkish breast imaging reports into structured and machine-readable datasets, with a specific focus on the extraction and standardization of BI-RADS labels already recorded in routine clinical reports. The proposed pipeline integrates domain-specific preprocessing, text normalization, hybrid conclusion-section segmentation, multi-phase BI-RADS label extraction, duplicate and near-duplicate report handling, and modality-based separation within a transparent rule-guided workflow. The framework was applied to two real-world breast imaging datasets comprising 35,104 reports in Dataset 1 and 27,215 reports in Dataset 2 after overlap and duplicate control around the May 2023 transition period. The datasets included ultrasonography, mammography, and magnetic resonance imaging records. The framework achieved conclusion-section segmentation coverage rates of 96.71% and 99.94%, respectively, and BI-RADS label extraction coverage rates of 95.04% and 100.00%, respectively. These coverage values indicate the proportion of reports for which the rule-guided pipeline produced extractable outputs and should not be interpreted as precision, recall, F1-score, or accuracy against an independent gold-standard corpus. The curation process reduced report-level textual redundancy, improved structural consistency, and enabled systematic separation of report narratives from diagnostic assessment labels. Manual review of selected subsets was used as an internal plausibility check rather than as a substitute for independent radiologist-annotated gold-standard validation. By addressing the challenges of structuring free-text breast imaging reports in a low-resource language setting, this study provides a transparent and adaptable methodological basis for future clinical NLP, machine learning, and real-world healthcare analytics.

S. Yildirim, Erkan Ülker, Necdet Poyraz · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.