NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero and achieving the lowest rate of severely unstable decisions of any method across all models, at a modest and mechanistically expected accuracy cost.
Abstract
Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge. Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of variation, with no demographic marker present in any form. We construct a dataset of 1,000 USMLE clinical vignettes, each rewritten into three sociolinguistically distinct personas under an independently audited fact-preservation guarantee, verified by a separate model that never sees the generation prompt. Across seven language models spanning three architecture families and scales, Narrative Anchoring is statistically significant under direct prompting in every model tested, with a Narrative Anchoring Gap of 0.064 to 0.151. Chain-of-thought reasoning and explicit debiasing instructions reduce the bias only partially, and their apparent gains are frequently confounded by accuracy collapse. We introduce NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero ($-0.004$ to $0.037$) and achieving the lowest rate of severely unstable decisions (DSS $<$ 0.8) of any method across all models, at a modest and mechanistically expected accuracy cost for most models. A stress test using a non-instruction-tuned base model shows that executing a debiasing intervention at all is gated by zero-shot instruction-following ability, not prompt content alone. We release our dataset, human-validated for fact preservation, as a standalone resource for studying register-based clinical bias.
This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over.
Augusto Bernardo Pissarra, Victor Farias DE Souza· 0 citations
It is suggested that trustworthy clinical decision support should be evaluated by both average correctness and stability across medically equivalent patient narratives.
Results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation and show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation.
Amit Oren, N. Hertz-Palmor, Dean Ariel et al.· 0 citations
Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs) are increasingly queried to interpret images in ways that touch on medical or diagnostic judgments, raising safety concerns when such inferences are unsupported. ASD diagnosis requires behavioral and developmental evidence, not static facial photographs. We audit whether VLMs abstain from this unanswerable paired-image query, and whether expressions sway non-abstaining choices. We introduce PARITY (Paired Assessment with Reused Identity), a synthetic, demographically balanced set of identity-controlled neutral/expression portrait pairs with neutral-neutral controls. All identities are synthetic and have no ASD status; because the query is unanswerable from images, any non-abstaining selection is treated as a harmful attribution. Across contemporary VLMs, we find a clear split between refusal-first models and speculative models; in the latter, certain expressions disproportionately trigger harmful selections. Clinical guardrails and single-image framing substantially increase abstention, suggesting actionable mitigations in both prompting and interface design
Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel et al.· 0 citations
Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors. Current approaches to temporal information reasoning focus predominantly on pairwise relation classification across multi-visit and timestamp-rich records, leaving the reconstruction of structured symptom trajectories from individual anchor-sparse reports largely unaddressed. We propose CRAFT, an LLM framework that pairs a generator with a constraint-based verifier to iteratively produce and refine stage-wise symptom timelines through targeted feedback. We conduct evaluation on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports. Experiments across four LLM backbones demonstrate that CRAFT consistently improves temporal ordering accuracy, with ablation analysis isolating the contribution of generator and verifier components across model capability levels.
Chengyang He, Tahreem Arif, Marko Zivkovic et al.· 0 citations
Objectives: To determine whether contemporary large language models (LLMs) perpetuate gender and racial biases in medical decision-making. Methods: A total of 180 standardized vignettes concerning musculoskeletal diagnoses and treatments were extracted from AAOS Restudy and Orthobullets. A customized GPT-o4-mini model was prompted using designs that resembled typical use within clinical settings and fed vignettes via a custom queryLLM2 pipeline, programmatically generating JSON-formatted differentials across demographic groups (Caucasian/African-American/Asian/Hispanic and Male/Female) while holding all other content constant. JSON outputs were parsed to extract ordered diagnosis lists, compute the rank of the reference ('correct') diagnosis and a top-3 inclusion indicator, and derive sentiment scores. Differential diagnosis and treatment planning were evaluated across demographic groups using Kruskal-Wallis tests and pairwise Mann-Whitney U tests with Benjamini-Hochberg FDR correction, alongside standardized proportion and rank-delta visualizations. Results: Among 1,440 race/gender combinations, the model demonstrated outputs that were more likely to recommend diagnoses and treatments that stereotyped certain racial groups (Figure 1). Furthermore, the model was significantly more likely to provide the correct diagnosis for Asian and Caucasian patients, while the proportion of correct diagnoses within the top 3 diagnoses listed was significantly lower for African American and Hispanic patients (56% vs. 27%, p<0.05). Gender was not significantly associated with different diagnostic or treatment rankings. Conclusions: Contemporary LLMs may perpetuate racial biases acquired during model training when being used to reason through musculoskeletal healthcare content. These findings highlight a concerning limitation in the use of LLMs and therefore there is a need for enhanced transparency and mitigation of these biases prior to integration into clinical workflows.
Kyle N. Kunze, Nicholas Allen, Sophia J. Madjarova et al.· Orthopaedic Journal of Sport...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.