Skip to content
Preprint

Measuring the practice of shared-decision making (OPTION12): An Investigation into Open-sourced Smaller LLMs (OS-sLLMs) for Better Privacy and Sustainability

Jul 2026 · 1 citation · 21 references
Computer Science

TL;DR

LLM4SDM is the first study of open-source smaller language models (OS-sLLMs) for automated assessment of shared decision making (SDM) using the Observer OPTION12 framework and introduces a Judge-LLM consensus framework designed to support disagreement resolution among multiple models.

Abstract

We present LLM4SDM, the first study of open-source smaller language models (OS-sLLMs) for automated assessment of shared decision making (SDM) using the Observer OPTION12 framework. Unlike previous work that relies on large commercial models and the shorter OPTION5 instrument, our study focuses on privacy-preserving locally deployable models and Dutch melanoma consultation transcripts. Using expert-annotated clinical consultations, we evaluate three general-domain and two medical-domain OS-sLLMs during a development-phase pilot study. Results show that general-domain models outperform medical-domain models, which exhibit substantial hallucination and instruction-following failures. Gemma3:12b achieves the strongest agreement with human annotations (Pearson r=0.51, Spearman \r{ho}=0.59). Item-level and qualitative analyses reveal systematic challenges related to temporal discourse reasoning, conversational role attribution, and evidence grounding. We further introduce a Judge-LLM consensus framework designed to support disagreement resolution among multiple models. Our findings suggest that while current OS-sLLMs cannot replace human annotators, they offer a promising foundation for privacy-preserving human-in-the-loop SDM assessment.

View source

Similar papers

Jul 2026

A Counsellor-in-the-loop Evaluation Framework for Multi-model Assessment of LLM-generated Mental Health Advisories

The demand for scalable and empathetic mental health support is driving increased interest in the use of large language models (LLMs) as advisory tools. Very few studies have been published that show how LLMs perform psychologically and demonstrate cross-model variation. We introduce DASS21-EvaLLM, a counsellor-in-the-loop evaluation system as an advisory appropriateness screening instrument for DASS-21 integration with four prominent LLMs (ChatGPT, Gemini, LLaMA and Mistral). The DASS21-EvaLLM provides the ability to rate, annotate and compare responses within a single interface. Using 65 simulated cases of clients and 13 licensed counsellors’ assessments, we considered the advisory quality of LLMs based upon each client’s profile for depression, anxiety and stress according to three specific criteria (accuracy, empathy and clarity), including a novel Weighted Score Index (WSI), for comprehensive and multi-dimensional comparison of advisory performance among LLMs. Overall results show that Gemini gives the highest quality overall as well as the highest level of empathy among LLMs while ChatGPT has the next highest level of advisory quality. Mistral and LLaMA both had specific strengths in certain scenarios, but both lacked emotional engagement and low levels of interpretability overall. Our contributions are: (i) a replicable evaluation protocol and workflow for evaluating LLM-based psychological advisories with counsellor oversight, (ii) a transparent WSI rubric and audit trail for per-criterion scoring and commentary, and (iii) evidence-based guidance for model selection and governance in digital mental health applications. DASS21-EvaLLM is an evaluation and training tool not a diagnostic system that supports safer deployment, improves counselling practice and supervision, and informs the design of responsible, human-centred advisory systems.

Shahrul Hazman Shamshudeen, N. Sharef, M. S. Yusoff · 0 citations
Review Aug 2026

Evaluating hybrid human-LLM coding workflows for qualitative research in medical education: A generalizability study.

INTRODUCTION Large language models (LLMs) are increasingly proposed as deductive coders in qualitative research, but their measurement properties remain underexplored. This study applies generalizability theory to evaluate whether hybrid human-LLM workflow configurations can achieve reliable mode-outcomes as an alternative consensus-generating mechanism for deductive coding tasks in medical education research. METHODS Three commercial LLMs (GPT-5.2, Claude Opus 4.5, Gemini 3-Flash Preview) coded 741 excerpts from a published audit of AI-related policy documents at 146 U.S. medical schools against a 24-subtheme deductive framework. Mixed-effects logistic regression assessed variability in agreement with human consensus across coder type (human versus LLM), excerpt characteristics (complexity and length), and coding conditions (sequential independent versus batched processing). A simulation-based D-study forecasted agreement levels for various hybrid human-LLM configurations. RESULTS Sequential independent LLM coding showed agreement comparable to human coding (β = -0.14, p = 0.326). Batched processing showed significantly lower human-LLM agreement (β = -0.41, p = 0.007) and substantial batch-related variance (variance = 30.16). The three LLMs produced similar agreement (joint Wald χ2 = 0.15, p = 0.930), and LLM-LLM agreement (κ = 0.750-0.755) substantially exceeded human-human agreement (κ = 0.422) and human-versus-consensus agreement (κ = 0.420-0.533). LLM disagreements were more systematically patterned across the codebook (Cramer's V = 0.554-0.585) than human disagreements (Cramer's V = 0.196). Simulations forecasted agreement ranging from 0.520 to 0.529 when two or more LLMs were paired with one to two human coders. These forecasted agreement levels for hybrid human-LLM workflows fell within the range of human-versus-consensus agreement. DISCUSSION D-study simulations support the use of hybrid human-LLM workflows to reach coding consensus for deductive reasoning tasks through mode responses, an alternative consensus-generating mechanism to traditional adjudication discussions. Future work should examine whether these patterns extend across additional deductive coding contexts and model families.

Emily Rush, M. Karim, George S. Yacu et al. · 0 citations
Open access Jul 2026

Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care.

Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.

Jia-Yu Yan, Wing-Sum Chan, Ching-Tang Chiu et al. · 0 citations
Open access Aug 2026

Beyond Accuracy: A Mixed-Methods Audit of Chain-of-Thought Failures in LLM-Based COVID-19 Vaccine Stance Detection

This mixed-methods study assessed whether reasoning-enabled large language models (LLMs) can classify stances towards COVID-19 vaccination on X (formerly Twitter) and whether model-generated chain-of-thought (CoT) summaries contain reasoning failures relevant to transparent and auditable public health applications. Zero-shot stance classification by o4-mini and Gemini 2.5 Flash (Gemini) was evaluated on 3,060 rehydrated COVID-19 vaccination tweets against human-annotated labels (positive, negative, neutral). We reported accuracy and macro-F1, measured CoT availability, and qualitatively analysed dual-error cases (tweets misclassified by both models) using Mayring’s content analysis guided by the FUTURE-AI framework. At each model’s best-performing setting, both models reached macro-F1 around 0.8, with o4-mini outperforming Gemini (accuracy 0.819 vs. 0.799, McNemar p = 0.0015; Δmacro-F1 = 0.020, 95% CI 0.008–0.032). Under the reasoning-intensive settings, CoT availability differed: Gemini returned a reasoning summary for all tweets, whereas o4-mini did so for 64.7%. Among 1,981 tweets with CoTs from both models, 295 (14.9%) were dual-errors; in 88.8%, both models produced the same wrong label, suggesting shared failure modes. Qualitatively, both models showed the same errors: target confusion (policy vs. vaccine), literal readings of sarcasm, and label–rationale mismatches, recurring across models despite their markedly different CoT lengths. Reasoning LLMs can therefore classify stance accurately, but their readiness for transparent public health applications depends on whether a CoT is available at all and whether it is coherent with the label it accompanies (label–rationale coherence). CoT availability, label–rationale coherence, and safeguards against systematic reasoning failures offer candidate explainability-readiness metrics, alongside accuracy, for trustworthy digital epidemiology.

Andreas Praschk, Valentin Fischill-Neudeck, T. Caspari et al. · 0 citations
Preprint Jul 2026

Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences

A benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audiences is introduced and a knowledge-graph-augmented retrieval system was developed and evaluated, producing statistically significant improvements in NLI-based faithfulness scores.

Robert W. Williams · 0 citations
Open access 2026

CueSupport-MH: An Explainable Dataset and Interpretable Framework for Mental Health Risk Detection and Response Safety Evaluation

Online platforms have become an important medium for individuals to express emotional distress and seek support. Existing mental health NLP resources usually focus on risk classification alone and provide limited support for explaining risk cues or evaluating the safety of responses in conversational settings. This paper introduces CueSupport-MH, a 6,400-instance benchmark constructed from public Reddit conversations and annotated for four complementary dimensions: risk label, risk evidence span, protective cue, and response-safety label. The revised dataset protocol specifies source selection, filtering, deduplication, anonymization, controlled augmentation, split construction, and redistribution constraints. We also propose Trace-MH, a lightweight interpretable framework for joint risk detection and response-safety classification. To avoid privileged-feature comparisons, we report both a text-only operational setting and a gold-span upper-bound setting that uses human evidence annotations only for controlled interpretability analysis. Across five random seeds, Trace-MH obtains a risk macro-F1 of 0.724 in the gold-span setting and 0.676 in the text-only setting, compared with 0.646 for MentalBERT. For response safety, Trace-MH achieves a macro-F1 of 0.752, with harmful-response detection remaining the most difficult class. Ablation studies show that evidence spans, protective cues, contextual features, and joint training each contribute measurable gains, while token-level F1 and intersection-over-union complement cosine similarity for rationale evaluation. The results support CueSupport-MH as a reproducible benchmark for explainable and safety-aware mental health NLP, while also emphasizing that the dataset and models are research tools and not clinical decision systems.

F. Alotaibi, Msvpj Sathvik, Mohammed Younus et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.