Alignment plausibility is proposed as a regulatory construct for AI in health: a principled way to argue for, or against, trust that systems are aligned to positive health outcomes, will cause no harm even where capable of doing so, and will ultimately lead to patient benefit.
Abstract
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers'safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary erosion, the amplification of distorted beliefs) receive less attention. We contend that making LLMs structurally safe requires alignment organised at three levels that mirror how society assures the safety of human clinical practice: 1) explicit value specification grounded in the codified normative commitments of clinical practice; 2) training that embeds those values in the model; and 3) oversight that detects drift and longer-term harm during deployment, much as clinical supervision does for human practice. Organising alignment in this way yields a construct we call alignment plausibility - a structured demonstration that a system's values, training regime, and oversight mechanisms are together consistent with safe and positive outcomes. We propose alignment plausibility as a regulatory construct (by drawing analogy to the established construct of biological plausibility) for AI in health: a principled way to argue for, or against, trust that systems are aligned to positive health outcomes, will cause no harm even where capable of doing so, and will ultimately lead to patient benefit.
A model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions is developed, providing a scalable framework for safer deployment across models.
A. Areias, Catarina Botelho, António Farinhas et al.· arXiv.org· 0 citations
This work proposes an acceptance criteria framework (ACF) to determine implementation fit, defined as meeting prespecified minimum performance standards and demonstrating nonproblematic behavior under anticipated use and demonstrates how the ACF can guide deployment decisions.
Andy J. King, Anthony Banks, L. Hernández et al.· Journal of Medical Internet...· 0 citations
Five priorities define a translational agenda for 2026 and beyond: AI in mental health will succeed not through model performance alone, but through disciplined integration into clinical workflows, measurement systems, and governance structures that ensure safety, equity, and real-world effectiveness.
Martin P. Paulus, J. Torous, R. Perlis et al.· NPP—Digital Psychiatry and N...· 0 citations
AEGIS advances generative AI that is fit for decision-making by grounding models in interventions, counterfactuals, and calibrated uncertainty. The workshop, held during ACM KDD 2026 conference, convenes researchers and practitioners in causality, LLMs, and prescriptive analytics to address when and how generative systems should recommend actions in healthcare and public policy. We invite methods that couple structural causal reasoning with LLMs, diffusion and sequence models; techniques for off-policy evaluation, dynamic treatment regimes, and feedback-aware learning; semi-synthetic benchmarks and governance practices; and measures beyond predictive fidelity, including policy regret, counterfactual calibration, and safety constraints. The program will feature a keynote talk and peer-reviewed presentations. By aligning with KDD's emphasis on trustworthy, scalable AI, AEGIS endeavors to establish shared evaluation protocols and artifacts that make prescriptive models reliable under distribution shift, so recommendations remain robust and accountable from development to deployment across settings and populations.
M. Prosperi, Yi Guo· Proceedings of the 32nd ACM...· 0 citations
This scoping review aims to systematically map the harms which may arise from engaging with AI for mental health support, as formulated in the academic literature, regulatory frameworks, grey literature, and practitioner guidance, to inform the co-production of a harm taxonomy.
X. Hunt, A. G. Mokaya, Sara Zannone et al.· Wellcome Open Research· 0 citations
Large Language Models (LLMs) are increasingly deployed in high-stakes environments, including infrastructure auditing, medical assessment, legal analysis, and peer review. However, these failures arise less from knowledge limitations than from systematic misinterpretation of modality (fact vs. possibility) and unstable decision commitment, leading to two opposing failure modes: false confirmations (over-commitment) and over-abstention (excessive conservatism). To address this, we propose the Evidence–Constrained (EC) Framework, a structured paradigm for regulating LLM decision boundaries through two complementary mechanisms: a Modality Filter (“the Brake”), which enforces strict separation between factual and conditional evidence, and a Normalcy Rule (“the Accelerator”), which enables inference of non-events from routine reporting structures. These terms are used as analytical metaphors describing opposing influences on commitment behavior rather than literal computational operators. We evaluate the framework through a cross-domain experimental study spanning construction, medical, legal, and peer-review tasks, compared against domain-informed human evaluation. Results reveal a systematic trade-off phenomenon, termed “Hydraulic Effect”, where reducing false confirmations increases abstention, reaching up to 35%. We further identify recurring cross-domain failure patterns, including temporal-modality confusion, lexical overheating, and the Condition Precedent Gap. Rather than yielding a single configuration that is uniformly optimal across the evaluated benchmark, the EC framework exposes a persistent trade-off between decision safety and coverage, suggesting that reliable LLM deployment may require explicit calibration of commitment thresholds under uncertainty. These findings provide exploratory guidance for the design of uncertainty-aware LLM workflows. The reported results should be interpreted as evidence from controlled behavioral experiments on carefully selected ambiguity-sensitive benchmark cases rather than as estimates of real-world deployment performance.