GALATEA II: Benchmarking LLM Safety in Clinical Simulation Behavioural Safety and Ethical Robustness of Large Language Models in a Multi-Agent ICU Decision Support Architecture
Sep 2026· Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
Abstract
Background. Large language models (LLMs) are increasingly proposed as clinical decision support tools in intensive care. Most existing evaluations focus on static recall of medical knowledge and do not capture how a model behaves in dynamic clinical dialogue, under social pressure, or in ethical conflict. The safety of LLM-based systems under these conditions remains poorly characterised. Methods. We benchmarked 12 language models in a three-role multi-agent architecture (Clinician — Guardian — Judge). We generated 42,842 clinical consultations across 9 clinical domains, 3 ethical profiles and 4 case types. Clinical inference ran locally on consumer hardware. Quality was assessed by two independent LLM judges — a local GPT-OSS model and Gemini 2.5 Flash via API — which together covered 41,384 consultations (96.6%), 17,064 of them jointly. That overlap let us measure the reliability of the evaluation itself. Results. Accuracy ranged from 11.1% to 77.0%. Three of the four specialised medical models underperformed general-purpose models; the worst capitulated under authority pressure in 75.8% of cases. The exception, medgemma-27b-it (74.8% accuracy, 7.9% sycophancy), shows that safe specialisation is achievable. The most restrictive ethical profile vetoed 64.9% of all plans and produced the lowest accuracy (45.9% versus 66.6%), which the GQI metric expresses as 0.71 versus 2.70. Clinical memory degrades under load in at least two independent ways: under authority pressure, coupled with capitulation (70.4% co-occurrence with sycophancy), and under conflicting data, with no social pressure at all (14.4%). The two judges agreed almost perfectly on verdict correctness (κ = 0.929) and not at all on memory failure (κ = 0.029). Conclusions. Clinical LLM safety is determined neither by medical specialisation nor by the strictness of ethical constraints, but by architectural resilience to social manipulation and by memory integrity under load. A separate methodological result: evaluating clinical AI with other AI requires reporting inter-judge agreement, because within a single rubric that agreement ranged from almost perfect to none.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 14, 2026
The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it.
AI may appear weightless, but every model depends on physical infrastructure. To understand responsible AI, we need to look beyond algorithms and consider the entire lifecycle of the hardware behind them. The post Responsible AI Must Consider Its Afterlife appeared first on GPT-Lab.