Semantic de-identification of burned-in PHI in DICOM medical images: a deep learning–NLP pipeline validated on clinical and phantom TMM datasets
A semantic de-identification pipeline integrating YOLOv11n-based text detection, domain-optimized EasyOCR, and a hybrid natural language processing (NLP) classification module combining regular expressions, keyword matching, and named entity recognition is proposed, confirming that the pipeline preserves quantitative pixel fidelity when applied to institutional and device identifiers embedded in phantom acquisitions.