Skip to content
Open access

Reasoning vs. conventional large language models for BI-RADS educational questions answering: a multi-model comparative evaluation

Yu-Xia Tang Yu-Ting Liu Meng-Xuan Liu Hao Ni Si-Qi Wang Shouju Wang
Aug 2026 · Frontiers in Medicine · 0 citations · 26 references

Abstract

To compare reasoning vs. conventional large language models (LLMs) in generating answers with guideline-aligned explanations for Breast Imaging Reporting and Data System (BI-RADS) educational questions. In this prospective study performed from February 6 to 12, 2025, 49 English-Chinese question pairs were extracted from BI-RADS Atlas Fifth Edition. Two reasoning LLMs (ChatGPT-o1, Deepseek-R1) and six conventional LLMs (Gemini2.0-Flash, Deepseek-V3, ChatGPT-4o, ChatGPT-3.5, Qwen-2.5, and WenXinYiYan-3.5) generated answers and explanations to the questions through structured prompts. Three radiologists specialized in breast imaging independently evaluated responses using a 5-point Likert scale, with reference to standard answers. The reasoning LLMs significantly outperformed conventional models (median [interquartile range (IQR)]: 3.7 [2.7–4.0] vs. 2.7 [2.0–3.7], P  < 0.001), with ChatGPT-o1 and Deepseek-R1 demonstrating peak performance. Both categories of LLMs exhibited significant score reductions in handling questions with multifaceted clinical scenarios (reasoning models: median 4.0 [2.7–4.3] vs. 2.7 [2.3–2.7], Δ median = −1.3, P  < 0.001; conventional models: 2.7 [2.0–3.7] vs. 2.3 [2.0–2.7], Δ median = −0.4, P  < 0.001). While question language showed no significant impact on reasoning LLMs (ChatGPT-o1 and Deepseek-R1), it affected some conventional models (ChatGPT-3.5, Deepseek-V3 and Gemini2.0-Flash). LLMs performance remained independent of question section and question type. Reasoning LLMs show significant potential for BI-RADS guideline explanation and education, but require specific optimization for complex clinical scenario instruction.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.