Review
Open access
Aug 2026
Robustness Gap of Large Language Models in Nephrology
Newer models may be more robust, but multiple-choice accuracy remains an incomplete measure of clinical reasoning robustness, as all evaluated LLMs showed a significant robustness gap after NOTA replacement.
A. Soejima, F. Kitano, D. Ichikawa et al.
· medRxiv · 0 citations