Oct 2026· Frontiers in Medicine· 0 citations· 38 references
Abstract
Large language models (LLMs) have emerged as a novel technical solution for intelligent prescription review. Nevertheless, comparative analyses of their performance remain limited. This study comprehensively evaluates five mainstream LLMs (ChatGPT, DeepSeek, Kimi, Qwen and Doubao) in prescription review, aiming to furnish data-based references for model selection, optimization and clinical deployment of intelligent prescription review systems.
This study employed a cross-sectional comparative design. The evaluation dataset was derived from the prescription review question bank of the Chinese Pharmacist Skills Competition. In total, 100 prescriptions from 10 clinical specialties were randomly sampled, encompassing 270 standardized medication assessment items. Official standard answers served as the gold standard. Quantitative comparisons were performed on overall detection rate, question type identification and adaptability across clinical specialties. Prescription-level repeated-measures analysis of variance (ANOVA) and Holm-adjusted pairwise paired
t
-tests were used to examine inter-model statistical differences.
The five models yielded item-level prescription detection rate ranging from 45.19% to 70.37% across 270 standardized medication assessment items. Their performance ranked from highest to lowest as follows: Kimi, DeepSeek, Qwen, Doubao and ChatGPT. Overall, the four Chinese domestic models included in this study outperformed the single tested international model (ChatGPT) in this China-focused prescription review task. Repeated-measures ANOVA confirmed statistically significant within-prescription differences in overall performance (F = 22.46,
P
< 0.001). While all models performed well at identifying inappropriate dosage forms and administration routes, they were less effective in detecting issues involving absent allergy test documentation and flawed result interpretation. Prescription reviews for the respiratory department yielded the best results, whereas detection rate was substantially lower for specialized clinical specialties such as oncology, pediatrics and pain management. Even the best-performing model presented a high false-negative rate and cannot yet satisfy the strict reliability criteria for clinical application in prescription review.
Performance varied markedly across the five LLMs in prescription review. Although these models can effectively identify routine and straightforward medication issues, they demonstrate prominent deficiencies in assessing complex prescriptions from specialist clinical specialties and detecting hidden medication risks. Currently, LLMs are only suitable for preliminary auxiliary screening and cannot substitute dedicated prescription review software or manual pharmacist evaluation.
This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.
P. Abrahamsson, O. Salo, Jussi Ronkainen et al.· arXiv.org· 727 citations· ⚡54
The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.
M. Pikkarainen, Jukka Haikara, O. Salo et al.· Empirical Software Engineeri...· 401 citations· ⚡48
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.
O. Salo, P. Abrahamsson· IET Software· 238 citations· ⚡9
Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· Journal of Systems and Softw...· 236 citations· ⚡13
The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.
P. Abrahamsson, Antti Hanhineva, H. Hulkko et al.· Conference on Object-Oriente...· 225 citations· ⚡18
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.