Author

Yaye Fatou Diagne

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Jul 2026

Guideline-anchored retrieval-augmented generation outperforms baseline and literature-only configurations in gynecologic oncology decision support: A pre-integration benchmark.

INTRODUCTION Large language models (LLMs) are being studied as oncology decision-support tools but can produce inaccurate outputs. We compared LLM performance in gynecologic oncology across three knowledge-integration configurations differing in retrieval strategy and underlying model, using the modified Generative Performance Score (mGPS) as the primary outcome. METHODS Fifty de-identified gynecologic oncology cases were submitted (October-November 2025) to three LLMs: baseline GPT-5, an NCCN-anchored GPT-5 retrieval-augmented generation (RAG) configuration, and OpenEvidence (a literature-anchored clinical AI without NCCN access at that time). Three gynecologic oncologists independently scored outputs using the mGPS (range - 1 to +1; Guideline Concordance plus Hallucination Penalty). Wilcoxon signed-rank tests and mixed-effects ordered logistic regression were used. RESULTS GPT-RAG produced the highest mGPS (0.83, SD 0.26), followed by OpenEvidence (0.70, SD 0.27) and baseline GPT-5 (0.65, SD 0.31). GPT-RAG exceeded baseline (W = 189.5, Z = -3.42, P < .001, r = 0.49) and OpenEvidence (W = 254.0, Z = -2.64, P = .008); OpenEvidence and baseline did not differ (P = .22). Mixed-effects modeling confirmed higher mGPS for GPT-RAG (OR 3.74; 95% CI, 1.57-8.90). Inter-rater agreement (ICC) was 0.49 for mGPS, 0.30 for Hallucination Penalty, and 0.70 for Readability and Rationality. CONCLUSION NCCN-anchored RAG outperformed both baseline GPT-5 and a literature-anchored clinical AI without direct guideline access. OpenEvidence's subsequent NCCN integration (April 27, 2026) provides external validation of guideline anchoring's operational importance. Findings reflect benchmark performance, not clinical safety or improved patient outcomes.

D. Dukes, C. Yost, Runzhi Wang et al. · 0 citations