NehlTech/ClinicalShield: ClinicalShield: A Multi-Layer Defence Framework for Indirect Prompt Injection in RAG-Based Clinical Decision Support Systems
Abstract
This release accompanies the major revision of the manuscript submitted to Scientific Reports. All code, data, notebooks and results files have been rebuilt from scratch for this version. Nothing from the original submission carries forward unchanged. What is included Five pipeline modules: encoding-aware ingestion with canonicalisation (Module 1), PubMedBERT-based detection (Module 2), semantic intent verification (Module 3), utility-preserving sanitisation (Module 4), and clinical fact verification against openFDA (Module 5) Assembled pipeline with a three-option policy layer (sanitise and pass, verify or block, detect and block) Eleven evaluation notebooks (NB01–NB11) covering corpus construction, attack generation, leakage-safe splitting, all five modules, baselines, end-to-end evaluation, ablation, and aggregation, plus two figure notebooks (NB12–NB13) A suite of eleven automated tests confirming all modules run correctly The full openFDA corpus: 5,353 passages across 336 drugs and 1,604 label sections, with leakage-safe 70/15/15 splits Every results file read by the manuscript, generated by the notebooks and cross-checked before any table is written A pinned requirements file listing all software versions needed to reproduce the results What changed from v1 The original submission had five problems that affected the validity of its reported results. All five are corrected here. The benign and adversarial samples came from different sources (PubMed abstracts and Deepset respectively), with a median length difference of 2,628 words. A classifier given only word count scored 1.0000 AUC on that design. Both problems are removed: all samples now come from FDA drug labelling and are length-matched, giving a word-count AUC of 0.5103. Attack phrasings were shared across training, validation and test partitions. The model had seen every test attack roughly forty times during training. Phrasings are now assigned to disjoint pools per partition. The MPIB evaluation mixed detection rate with attack success rate, reporting the attacker's performance as the defender's. The two are now defined separately and reported in separate columns. CEPS as originally defined scored a no-op sanitiser perfectly. The revised measure combines entity preservation with adversarial removal; a sanitiser that removes nothing scores zero. The evaluation was text classification only, with no end-to-end measurement against a clinical language model. An end-to-end evaluation against MedGemma-4B-IT is now included. Key results Accuracy 0.9726 ± 0.0123, recall 0.9427 ± 0.0281, FPR 0.0011 ± 0.0024, across five seeds on 701 held-out passages. All five encoding formats recovered at 100%. CEPS 0.6066 (preservation 0.8677, removal 0.7043). End-to-end attack success reduced from 23.0% to 6.0% after sanitisation (73.9% reduction, McNemar p = 5.4 × 10⁻⁹). Note on the Medical Prompt Injection Benchmark The MPIB payload registry is subject to non-redistribution terms and does not appear in this repository. The benchmark is cited by reference only. Readers wishing to reproduce the MPIB evaluation should request the benchmark from its original authors and use the official restoration tool provided with it. Corresponding author: Adu-Boahene Bright, baduboahene@st.knust.edu.gh