Jul 2026· McMaster University Medical Journal· Vol 22, pp. 21· 0 citations
TL;DR
Preliminary findings demonstrate a feasible, scalable, and reproducible workflow for comparative LLM evaluation in fever in the returning traveller assessment.
Abstract
Background: Fever in the returning traveller is a common but challenging presentation with a broad, geography-dependent differential diagnosis. Timely assessment can be difficult for front-line clinicians, especially outside tropical medicine settings. Large language models (LLMs) may support clinical decision-making by generating differential diagnoses, but comparative evaluation workflows remain underdeveloped.
Objective: To develop and evaluate an LLM workflow to assess the reliability and efficacy of LLM-generated differential diagnoses for written fever-in-returning-traveller cases, compare performance across four LLMs, and secondarily evaluate LLM case-generation capability.
Methods: We studied 21 travel-related diagnoses using four LLMs (ChatGPT 5.2 Thinking, GPT-o3, Llama 3.2, and Mistral Instruct). Infectious diseases fellows created cases across the diagnoses (n=84; 4/diagnosis), which were used to calibrate AI case-generation prompting. Each LLM then generated matched cases (n=84; 21/model). Clinician and AI-generated cases were pooled, assigned unique IDs, and randomized so each analysis model received equal numbers of clinician and AI cases (n=168 analyses; 42/model). Feasibility outcomes were completion, retries/system errors, and logged response time. Expert qualitative scoring is underway.
Results: We completed dataset generation and model-output acquisition for 252 outputs across 21 diagnoses. Case-generation retries occurred in 17/84 outputs (20.2%), all first-pass refusals were from Llama 3.2, all resolved with repeat prompting. Case-analysis retries were rare (1/168, 0.6%; single no-response, resolved), with no persistent failures. Balanced allocation was achieved. Median case-analysis response times for cloud-based models were 11s (ChatGPT 5.2 Thinking) and 4s (GPT-o3).
Conclusions: Preliminary findings demonstrate a feasible, scalable, and reproducible workflow for comparative LLM evaluation in fever in the returning traveller assessment. Model-specific refusal behaviour is an important implementation consideration. Ongoing expert qualitative scoring will determine comparative reliability and efficacy. Current conclusions are limited to operational feasibility and dataset generation reliability.
On complex ID scenarios, large language models responses were variable and caution is required when deploying these models in ID domains without specialist oversight, suggesting caution is required when deploying these models in ID domains without specialist oversight.
A. Pradhan, B. Waxse, W. Matias et al.· medRxiv· 0 citations
A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.
Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al.· Indian Journal of Computer S...· 0 citations
Large language models (LLMs) are emerging as tools to support clinical decision making. HIV management is a compelling use case due to its complexity and dynamic nature, involving diverse treatment options, comorbidities, and adherence challenges. However, integrating LLMs into clinical practice raises concerns about accuracy, safety, and clinician acceptance. Despite growing interest, their performance in HIV care remains poorly studied, and benchmarking is lacking.
We developed HIVMedQA, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios. We evaluated seven general-purpose and three medical LLMs. Performance was assessed using lexical similarity and an extended medical LLM-as-a-judge framework capturing key clinical dimensions, including question comprehension, reasoning, knowledge recall, bias, potential harm, and factual accuracy, to better capture nuances relevant to the medical domain, with additional evaluation by HIV-experienced physicians.
Performance varies substantially across models and task complexity. Gemini 2.5 Pro achieves the highest overall scores, followed by Claude 3.5 Sonnet and MedGemma-27B. Knowledge recall is generally stronger than question comprehension or clinical reasoning. Medical LLMs do not consistently outperform general-purpose models, and model size alone does not predict performance. Several models are sensitive to cognitive bias prompts. LLM-as-a-judge scoring aligns better with clinician assessment than lexical metrics.
HIVMedQA provides a structured benchmark for evaluating LLMs in HIV clinical decision support. Current LLMs show promise, but limitations in reasoning, bias robustness, and safety indicate that careful validation, domain-specific evaluation, and clinician oversight remain essential before clinical deployment.
Gonzalo Cardenal-Antolin, J. Fellay, Bashkim Jaha et al.· Communications Medicine· 0 citations
Current evidence supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
Yun-Jia Wu, Qi Yan, Dingcheng Tian· AI Medicine· 0 citations
In this scenario-based early evaluation, GPT-5 was the most safety-aligned model, GPT-4 was the most operationally actionable, and Gemini was strongest for synthesis, but all require clinician oversight, prospective validation, and governance before clinical deployment.
Y. K. Çalışkan, Fatih Başak, Olgun Erdem· World Journal of Surgery· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.