Skip to content
Open access

General-Purpose vs. Domain-Specific Large Language Models in Antibiotic Clinical Decision-Making: A Double-Blind Evaluation with a 2X2 Factorial Design

Jul 2026 · medRxiv · 0 citations
Medicine

TL;DR

Domain-specific medical LLMs enhanced by CoT approach the antibiotic decision-making level of real physicians, with advantages in individualization and dosing precision, but notable deficiencies persist in antimicrobial stewardship ecological awareness and automated evaluation reliability, underscoring the continued indispensability of senior clinical expertise.

Abstract

Background: Antimicrobial resistance poses a major threat to global public health. Large language models (LLMs) offer new possibilities for optimizing antibiotic prescribing decisions, but the capabilities of general-purpose versus domain-specific medical LLMs under different prompting strategies remain to be clarified. Methods: This double-blind, randomized-sequence evaluation used a 2X2 factorial design comparing four AI conditions-the domain-specific model MedGo and the general-purpose model DeepSeek V3.5, each under standard direct prompting and chain-of-thought (CoT) prompting-alongside real physician prescriptions across 59 complex inpatient infection cases. Five parallel regimens were generated per case and independently evaluated by three senior clinicians (1-5 comprehensive score and five domain sub-scores). ChatGPT 5.2 was additionally assessed as an automated evaluation tool. Results: Score ranking: real physicians > MedGo-CoT > DeepSeek-CoT > MedGo> DeepSeek (Friedman test, p<0.001). In base mode, MedGo significantly outperformed DeepSeek (Holm-adjusted p=0.040). CoT improved both models (Holm-adjusted p<0.001 for DeepSeek; p=0.024 for MedGo) and reduced score dispersion. MedGo-CoT significantly outperformed DeepSeek-CoT in individualized adjustment (adjusted p<0.001) and dosing precision (adjusted p=0.005). ChatGPT-expert correlation was negligible (overall Kendall {tau}=0.153, p=0.003; subgroup {tau}=0.06-0.20, all p>0.05). Conclusions: Domain-specific medical LLMs enhanced by CoT approach the antibiotic decision-making level of real physicians, with advantages in individualization and dosing precision. However, notable deficiencies persist in antimicrobial stewardship ecological awareness and automated evaluation reliability, underscoring the continued indispensability of senior clinical expertise.

Read PDF

Similar papers

Open access Jul 2026

Evaluation of four large language models on complex, infectious disease case scenarios

On complex ID scenarios, large language models responses were variable and caution is required when deploying these models in ID domains without specialist oversight, suggesting caution is required when deploying these models in ID domains without specialist oversight.

A. Pradhan, B. Waxse, W. Matias et al. · 0 citations
Open access Aug 2026

Decoding high-order clinical correlations: a knowledge-driven large language model framework for specialized medical decision-making

Embedding domain-specific knowledge into LLMs may improve performance on specialized exam-style thoracic-surgery questions on this text-only benchmark, however, the present 56-item evaluation does not establish clinical equivalence, diagnostic accuracy in practice, multimodal competence, or readiness for real-world clinical decision support.

Qian Li, Yongxin Li, Chao Ye et al. · 0 citations
Preprint Aug 2026

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

Praveen Reddy, C. Mandke, Suvrankar Datta et al. · 0 citations
Open access Jul 2026

Artificial Intelligence and Large Language Models as Decision-Support Tools in Hospital Compounding Pharmacy: A Proof-of-Concept Study.

Objectives To evaluate the efficiency and reliability of a large language model (LLM) as a decision-support tool in hospital compounding pharmacy for pediatric extemporaneous preparations requiring assessment of drug crushability, regulatory compliance, and formulation feasibility. Methods A proof-of-concept study compared a structured LLM-assisted workflow with the traditional manual information retrieval process in a hospital pharmacy setting. The LLM (Claude Sonnet 4.6) was configured with a standardized prompt to extract and consolidate data from multiple authoritative sources: the Friuli Venezia Giulia "Do Not Crush" list, the Italian Medicines Agency (AIFA) database for Summary of Product Characteristics (SmPC), AIFA Law 648/96 off-label use lists, and the Stabilis database for oral liquid formulation stability. Two representative drugs, propranolol hydrochloride and imatinib mesylate, were analyzed. For each drug, the model generated a structured output including crushability, regulatory information, off-label status, and extemporaneous formulation data. The same queries were manually performed by an experienced hospital pharmacist. Primary outcome was information retrieval time; secondary outcomes included completeness and accuracy. Results The LLM-assisted workflow reduced retrieval time to less than 2 minutes per drug (mean 1 minute 45 seconds), compared with a mean of 20 minutes (range 15-25 minutes) for the manual process, plus an additional 5 to 10 minutes for transcription. Output completeness was 100%, with all predefined fields correctly populated. The model accurately classified drug crushability and correctly identified Law 648/96 regulatory status. For propranolol, the system identified crushability, pediatric off-label authorization, and SyrSpend-based formulations with stability data of up to 146 days at room temperature. For imatinib, the model highlighted cytotoxic handling precautions, identified the absence of SyrSpend formulations, and retrieved alternative formulation stability data (30 days refrigerated). Conclusions A properly configured LLM can function as an effective decision-support tool in hospital compounding pharmacy, improving efficiency while maintaining high standards of completeness, accuracy, and regulatory compliance. These preliminary results support further investigation into the integration of the LLM system into routine pediatric galenical preparation practice.

E. Castellana, MR Chiappetta · 0 citations
Review Open access Aug 2026

Benchmarking publicly accessible large language models for English-language patient-facing acute pancreatitis information: a cross-sectional study of quality, transparency, and readability

Background Patients increasingly rely on large language models (LLMs) for health information, yet their suitability for decision-critical conditions such as acute pancreatitis remains unclear. Given that acute pancreatitis requires timely symptom recognition, severity assessment, treatment decision-making, recurrence prevention, and follow-up management, LLM-generated information should demonstrate reliability, transparency, and readability. Objectives To evaluate the informational quality, visible transparency-related features, and readability of English-language responses generated by five publicly accessible LLMs to standardized patient-facing questions on acute pancreatitis. Methods This cross-sectional benchmark study developed 24 English-language, single-intent questions on acute pancreatitis across six clinical domains using public search intents and guideline-derived decision-critical content. The analysis focused exclusively on English-language patient-facing responses. Each question was submitted once to GPT-5.4 Thinking, DeepSeek-V3.2, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview, generating 120 responses. Anonymized responses were independently evaluated by two blinded gastroenterologists using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), and Journal of the American Medical Association (JAMA) benchmark criteria. Readability was assessed using six established formulas. A structured response-level safety analysis was added to evaluate factual inaccuracies, clinical hallucinations, and clinical safety-risk severity. Between-model differences were analyzed with Friedman tests followed by post hoc paired Wilcoxon signed-rank tests with Holm correction. Results Significant between-model differences were observed across all quality, transparency-related, and readability outcomes. Grok 4.3 achieved the highest mean DISCERN, GQS, EQIP, and JAMA scores, reflecting the strongest informational quality profile according to the predefined quality instruments, although visible transparency cues remained limited across all models. In the added safety analysis, factual inaccuracies and clinical hallucinations were each identified in 16 of 120 responses (13.3%), whereas clinical safety-risk signals were identified in 7 responses (5.8%), all of which were adjudicated as score 1 (low risk) under the predefined clinician-rated rubric; no moderate- or high-risk clinical safety event was adjudicated. DeepSeek-V3.2 demonstrated the most favorable readability profile, with the highest Flesch Reading Ease score (44.42 ± 12.27), which nevertheless remained substantially below the recommended threshold of ≥ 80. None of the 120 responses satisfied all six predefined readability targets. All model-level readability distributions differed significantly from recommended thresholds in the direction of poorer readability. Conclusion Publicly accessible LLMs generated English-language responses on acute pancreatitis with variable informational quality, limited visible transparency cues, and consistently inadequate readability under standardized default public-interface conditions. Because the primary analysis used a single-generation cross-sectional design, model rankings should be interpreted as performance snapshots rather than definitive or temporally stable hierarchies. Higher quality scores did not necessarily establish factual accuracy, clinical safety, patient-education-level readability, or stronger response-level transparency. Current public-interface LLM outputs may serve as clinician-reviewed drafts for patient education but should not function as standalone patient resources. Future AI-based health information systems should strengthen clinical completeness, plain-language communication, visible evidence support, actionability, and clinician-supervised safeguards.

Biao Jiang, Hong-Xin Sun, Linlin Chen · 0 citations
Open access Aug 2026

A Human-in-the-Loop Large Language Model System Based on the Model Context Protocol for Differential Diagnosis from Electronic Medical Records and Literature

DDx-Finder is presented, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns.

H. Lim, H. Yi, J. Yoon et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.