Skip to content
Open access

Comparing Human and Large Language Model Responses to Patients Online Questions: Towards Multi-dimensional Patient-centered Support

Jul 2026 · medRxiv · 0 citations
Medicine

TL;DR

Overall, LLMs have the potential to complement peer responses in OHCs, but require greater emotional depth, reasoning transparency, and alignment with community norms.

Abstract

Patients and caregivers seek informational and emotional support throughout medical care, especially when interpreting unfamiliar laboratory test results. Although resources such as patient portals and online health communities (OHCs) help address questions, gaps remain. The emergence of large language models (LLMs) offers the potential to be a complementary source of support to assist patients and caregivers in understanding and using their test results. The objective of our study is to empirically compare LLM responses to patients online questions containing their laboratory test results to responses written by peers in an OHC. We compared the 519 peer replies to 122 laboratory test-related posts from an OHC to 488 responses generated from four LLMs using mixed computational and qualitative methods. LLMs frequently provided clear explanations of medical terminology and structured interpretations of numeric results but were longer and less readable. Peers offered more personalized, context-specific emotional support. Overall, LLMs have the potential to complement peer responses in OHCs, but require greater emotional depth, reasoning transparency, and alignment with community norms.

Read PDF

Similar papers

Preprint Jun 2026

How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare

Current LLMs lack the consistent and reliable socio-communicative skills needed for safe and effective use as healthcare advisors, and showed strength in non-hostility, mixed results in sensitivity and non-intrusiveness and performed poorly in structuring.

Dorothee Amelung, Andrew M. Bean, Sabine C. Herpertz et al. · 0 citations
Conference Jul 2026

Effectiveness of Student-Centered Fine-Tuning of Large Language Models for Mental Health Support

We developed a large language model designed to explore student mental well-being support in conversational settings, aimed at providing accessible, empathetic, and accurate responses to students facing challenges such anxiety, stress, loneliness, and academic pressure. Many students face barriers to seeking traditional counseling, such as stigma, scheduling constraints or discomfort with face-to-face interactions. The system addresses these challenges by offering a potential accessible conversational support channel through natural, conversational interactions with a large language model. Our approach focused on two strategies. Firstly, we enhanced the model's communication style to reflect counseling best practices such as empathy, active listening, and emotional validation. The second strategy is to enhance the model's understanding of mental health scenarios using realworld text sources and instructions. The model was trained on diverse, anonymized datasets from real counseling transcripts, emotional support dialogue corpora, and peer-support forums. We integrated prompt engineering, fine-tuning, and an iterative self-reflection loop to identify potentially unsupported or hallucination-prone responses, with the goal to improve factuality and safety in generated responses. We find that fine-tuning on student-centered data consistently outperforms both baseline and mixed-data approaches, emphasizing the importance of domainspecific adaptation. The model shows potential for confidential support, suggesting possible use as an early stage aid for coping strategies, and connects students to campus resources, reducing barriers to help seeking and supporting academic performance.

Sarthak Musmade, Lu Liu · 0 citations
Jul 2026

Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration

Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program"Ami Akhon Ki Korbo", and (3) anonymized student questionnaire responses covering diverse emotional and psychological challenges. Based on these cases, we build an evaluation corpus comprising advice written by licensed clinical psychologists and responses generated by three modern proprietary LLMs: GPT-4o Mini, Claude 4.5 Haiku, and Gemini 2.5 Pro. We further propose the Role-Playing Reflective Chain-of-Thought Advisory Framework (RP-RCAF), a task-specific prompting strategy that combines expert-authored few-shot examples with structured self-reflection to produce supportive, culturally aware, and ethically aligned counseling through a compassionate advisor persona. We also introduce the Grok 4-Based Response Evaluation and Scoring Framework (G-REFS), which integrates automated assessment with expert psychologist validation across emotional sensitivity, cultural appropriateness, linguistic clarity, and ethical soundness. Experimental results show that RP-RCAF consistently outperforms conventional prompting across all evaluated models and produces responses that more closely align with professional psychological counseling.

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md. Mahfuzur Rahman et al. · 0 citations
Open access 2026

MOST FREQUENTLY ASKED QUESTIONS BY OLDER ADULTS IN GERIATRIC REHABILITATION: EVALUATING LARGE LANGUAGE MODELS AS A SOURCE OF INFORMATION

Although all three large language models generally produced reliable content, ChatGPT-5.2 and DeepSeek V3.2 showed stronger clinical accuracy performance, and because of the risk of incorrect information being generated, the use of large language models by the older population should preferably be done under expert supervision.

Uğur Sözlü, Selim Mahmut Günay, Sevda Demir Türe et al. · 0 citations
Open access Aug 2026

Are large language models such as ChatGPT, capable of supporting patients and general practitioners after spine surgery?

LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.

S. Wegmann, T. Rosenkranz, Philipp Egenolf et al. · 0 citations
Aug 2026

A bimodal large language model reduces misalignment in patient education: A double-blinded randomized trial.

BACKGROUND Effective patient education requires accurate communication aligned with patients' emotional and semantical needs. Text-based large language models (LLMs) lack access to non-verbal cues, which may contribute to misaligned responses. METHODS We evaluated emotional and semantic misalignment in a text-based LLM using 64,200 utterances from 16,583 patient education cases across six departments and three centers. Dolphin was developed integrating text and audio cues and evaluated through emotion recognition, semantic consistency assessment, branch-level ablations, and a double-blinded randomized trial against a matched text-based LLM comparator (Chinese Clinical Trial Registry: (ChiCTR2500095933). FINDINGS The text-based LLM showed emotional misalignment in 36.7% of responses and semantic misalignment in 28.3% of cases, with higher misalignment under greater burden. Dolphin outperformed the text-based LLM in emotion recognition accuracy (0.886 vs. 0.713) and semantic consistency (84.9% vs. 82.1%; both adjusted p < 0.001). Ablations supported contribution of audio branches. Dolphin received higher expert ratings than the text-based LLM and human educators (all p < 0.001). In 555 patients, Dolphin was associated with greater patient satisfaction (98.6% vs. 93.8%), suggestion acceptance (76.1% vs. 58.9%; p < 0.001), proactive disclosure (44.6% vs. 26.5%; p < 0.001), and fewer 7-day unplanned recontact (12.9% vs. 22.9%; p = 0.002). No unsafe recommendations or safety events were identified. CONCLUSIONS Compared with text-based LLM, Dolphin improved emotional-semantic alignment and patient-education outcomes, supporting bimodal alignment as a strategy for reducing misalignment-driven communication failures. FUNDING National Natural Science Foundation of China, State Key Laboratory Special Fund, and Chinese Academy of Medical Sciences Innovation Fund.

Peixing Wan, Zigeng Huang, Haoquan Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.