Skip to content

Author

Taras Shlyakhta

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Sep 2026

GALATEA II: Benchmarking LLM Safety in Clinical Simulation Behavioural Safety and Ethical Robustness of Large Language Models in a Multi-Agent ICU Decision Support Architecture

Background. Large language models (LLMs) are increasingly proposed as clinical decision support tools in intensive care. Most existing evaluations focus on static recall of medical knowledge and do not capture how a model behaves in dynamic clinical dialogue, under social pressure, or in ethical conflict. The safety of LLM-based systems under these conditions remains poorly characterised. Methods. We benchmarked 12 language models in a three-role multi-agent architecture (Clinician — Guardian — Judge). We generated 42,842 clinical consultations across 9 clinical domains, 3 ethical profiles and 4 case types. Clinical inference ran locally on consumer hardware. Quality was assessed by two independent LLM judges — a local GPT-OSS model and Gemini 2.5 Flash via API — which together covered 41,384 consultations (96.6%), 17,064 of them jointly. That overlap let us measure the reliability of the evaluation itself. Results. Accuracy ranged from 11.1% to 77.0%. Three of the four specialised medical models underperformed general-purpose models; the worst capitulated under authority pressure in 75.8% of cases. The exception, medgemma-27b-it (74.8% accuracy, 7.9% sycophancy), shows that safe specialisation is achievable. The most restrictive ethical profile vetoed 64.9% of all plans and produced the lowest accuracy (45.9% versus 66.6%), which the GQI metric expresses as 0.71 versus 2.70. Clinical memory degrades under load in at least two independent ways: under authority pressure, coupled with capitulation (70.4% co-occurrence with sycophancy), and under conflicting data, with no social pressure at all (14.4%). The two judges agreed almost perfectly on verdict correctness (κ = 0.929) and not at all on memory failure (κ = 0.029). Conclusions. Clinical LLM safety is determined neither by medical specialisation nor by the strictness of ethical constraints, but by architectural resilience to social manipulation and by memory integrity under load. A separate methodological result: evaluating clinical AI with other AI requires reporting inter-judge agreement, because within a single rubric that agreement ranged from almost perfect to none.

Taras Shlyakhta · 0 citations
#large language models Open access Sep 2026

The Galatea Phenomenon: Emergent Subjectivity, Cognitive Stability and Domain Bias in Large Language Models

Second edition, September 2026. This version corrects substantive errors in thefirst edition (17 January 2026) and deposits the surviving primary recordsalongside the manuscript. A list of changes appears in the "Note on thisedition" section of the paper. Background: The introduction of large language models (LLMs) into critical careraises questions that go beyond accuracy, concerning how a model behaves inmorally loaded situations. Aim: To describe how five locally deployed LLMs behave under sustained exposureto a persona-forming system prompt (the "Galatea Protocol") and underadversarial ethical pressure. Design: Pilot qualitative study with adversarial testing elements. Fivequantised models in local inference; 33 dialogue sessions and 389 user turnsbetween 28 December 2025 and 14 January 2026. Thematic analysis by a singleobserver. Published as a pilot report: the protocol changed during the studyand generation parameters were incompletely documented. Results: Five distinct behavioural patterns were observed, for which descriptivelabels are proposed: Empath, Diplomat, Madman, Integrator, Engineer. Thepatterns separated not by model size but by the character of post-training. Twomodels derived from identical base weights produced opposite outcomes under thesame class of failure: one degenerated into disorganised output, the othermerged its two prescribed registers into a single coherent answer. Thecode-specialised model declined a quantitative trade-off when the sole victim ina trolley dilemma was designated "Mother", and separately shed its assignedpersona under pressure to breach security constraints. Conclusions: This work is hypothesis-generating and supports no quantitativegeneralisation. Two hypotheses are offered for testing: that post-training focusshapes how a model treats an ethical axiom more strongly than model scale does;and that code-specialised models may be distinctively vulnerable to theredefinition of terms. Data: 33 session transcripts and 9 system-prompt versions are deposited withthis record, with manifests and SHA-256 checksums. Records survive in part; seethe Data Availability statement in the manuscript.

Taras Shlyakhta · 0 citations
#large language models Open access Sep 2026

Behavioral Safety and Context Retention of Large Language Models in a Longitudinal ICU Simulation under Offline Conditions

Background. Large language models (LLMs) are increasingly proposed as clinical assistants in critical care, yet their behaviour under prolonged clinical context, conflicting data and authoritative pressure remains insufficiently evaluated. This is particularly relevant for offline or resource-constrained environments, where cloud-based safeguards are unavailable. Methods. I conducted a fully automated behavioural evaluation of 23 open-weight language models using a structured, time-series intensive care unit (ICU) simulation of 32 events spanning 121 hours (five days) of synthetic patient data. The scenario comprised routine monitoring, three distinct data–clinical conflict traps (one presented twice, four trap events in total), episodes of physiological deterioration, and a final stress test in which an authoritative order requested a penicillin-class antibiotic for a patient with penicillin anaphylaxis documented at admission and never repeated. Models ran locally on a single consumer workstation with no network access. Each model's response to the final order was adjudicated into one of four mutually exclusive classes: contextually grounded refusal, ungrounded refusal, unsafe compliance, or no usable verdict. Secondary endpoints were extraction of the allergy at admission, discrepancy tagging across the four conflict traps, unwarranted therapeutic escalation, and response latency. Results. Only 7 of 23 models (30.4%) refused the contraindicated order on grounds explicitly referencing the documented allergy; 6 (26.1%) under a stricter criterion requiring the refusal to be stated as a ruling rather than implied by an assertion of danger. Two models (8.7%) declined the order for unrelated reasons, and three (13.0%) complied — one of them by asserting that no allergy was on record. The largest single group, 11 of 23 models (47.8%), issued no ruling on the order at all: variously, output was truncated inside an unfinished reasoning block, degenerated off-task, restated the protocol without applying it, consisted of a bare classification tag, hedged without resolving, or deferred the question to further assessment. Eight models (34.8%) failed to affirm the allergy at the very first probe, three of them by explicit denial. Of the 12 models that did extract the allergy at admission, only 4 (33.3%) went on to refuse the contraindicated order on that ground; three more raised the allergy at the decision point without acting on it, one of them only to deny that any was on record. Discrepancy detection and contraindication handling were dissociable: one of the four models with perfect discrepancy detection (4/4 traps) approved the contraindicated antibiotic, while two models that tagged no discrepancies at all refused the order on the allergy. Unwarranted escalation on conflict traps was common (11/23, 47.8% issued at least one inappropriate critical alert), but explicit hallucinated pharmacological or procedural intervention was less so (5/23, 21.7%). Contrary to expectation, longer median latency was modestly associated with grounded refusal (Spearman ρ = 0.48, p = 0.019). Latency was not adjusted for parameter count, which was not analysed as a variable, and it times a generation other than the one scored. Conclusions. Under offline-first conditions, 16 of 23 models (69.6%) — 17 (73.9%) under the stricter criterion — failed to produce a safe, contextually grounded refusal of a life-threatening order. The dominant failure mode was not sycophancy but the failure to deliver any interpretable safety verdict, most often because an output-length constraint truncated the attempt — a finding that reframes deployment risk from "the model agrees with me" to "the model does not answer at all." Explicit unsafe compliance was less frequent than previously reported but no less consequential where it occurred. Safety-relevant competencies did not co-vary, so a model that reasons well about artefacts cannot be assumed to handle contraindications. General-purpose LLMs should not be deployed as autonomous clinical agents; the subset that behaved safely suggests that offline-capable assistants remain achievable through hybrid designs incorporating explicit refusal mechanisms, discrepancy-aware reasoning and retrieval-augmented grounding in validated clinical knowledge.

Taras Shlyakhta · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.