A Multi-Domain Human Expert Evaluation of Clinical and Behavioral Knowledge in Large Language Models
Abstract
Large Language Models (LLMs) such as ChatGPT and Gemini are increasingly used to answer medical and psychological questions, yet systematic evaluations across domains with differing reasoning demands remain limited. We present a multi-domain expert evaluation of two state-of-the-art models, ChatGPT Pro (v5.2) and Gemini 3 Pro, across three healthcare domains: Gynecology, Pathology, and Psychology. We curated 300 open-ended, realistic questions, 100 per domain, designed to elicit clinical reasoning, mechanistic interpretation, and conceptual explanation. Responses were independently scored by domain experts using a standardized five-point rubric. Results reveal domain-dependent performance patterns, with both models performing well in guideline-aligned advisory tasks in Gynecology, lower in mechanistic diagnostic contexts in Pathology, and showing divergent strengths in conceptual psychology questions. Quantitative and qualitative analyses highlight recurring limitations in contextual nuance, mechanistic depth, and safety framing. These findings underscore the importance of domainstratified evaluation and expert oversight when deploying LLMs for healthcare information and provide a reproducible framework for future assessments.