Aug 2026· Kaohsiung Journal of Medical Sciences· pp.
e70271
· 0 citations
Medicine
TL;DR
It is suggested that primary spoken language may be associated with ANT1 performance in MHE assessments and Integrating ANT1 with serum IL-6 showed numerically improved discrimination in Mandarin speakers, whereas exploratory demographic calibration of S-ANT1 showed a numerically higher AUROC in Taiwanese Hokkien speakers.
Abstract
Minimal hepatic encephalopathy (MHE) involves subtle cognitive dysfunction and systemic inflammation and is associated with an increased risk of overt hepatic encephalopathy. The Animal Naming Test (ANT1) is a rapid semantic fluency tool for MHE assessment, but its performance across different primary spoken languages remains unclear. We conducted a prospective proof-of-concept study to evaluate the diagnostic performance of ANT1 and serum interleukin-6 (IL-6) in Mandarin- and Taiwanese Hokkien-speaking cirrhotic patients. A total of 65 cirrhotic patients and 34 healthy controls were enrolled. Patients completed ANT1, simplified ANT1 (S-ANT1), and standard psychometric assessments. MHE was defined by abnormal PHES and/or visually assessed EEG slowing. Diagnostic discrimination was evaluated using AUROC analyses stratified by primary spoken language. A post hoc exploratory Taiwanese-calibrated S-ANT1 was also assessed in Taiwanese Hokkien-speaking patients. Sixteen cirrhotic patients (24.6%) were diagnosed with MHE. Patients with MHE had lower ANT1 scores and higher serum IL-6 levels than those without MHE. In Mandarin-speaking patients (n = 44), ANT1 demonstrated an AUROC of 0.760, while serum IL-6 showed an AUROC of 0.841. The composite model combining ANT1 and serum IL-6 showed a numerically higher AUROC of 0.895, but this improvement was not statistically significant in pairwise DeLong comparisons. In Taiwanese Hokkien-speaking patients (n = 21), standard ANT1 showed a lower AUROC of 0.679. The exploratory Taiwanese-calibrated S-ANT1 showed a numerically higher AUROC of 0.776; however, this post hoc finding was based on a small subgroup with 7 MHE events and requires external validation. This study suggests that primary spoken language may be associated with ANT1 performance in MHE assessments. Integrating ANT1 with serum IL-6 showed numerically improved discrimination in Mandarin speakers, whereas exploratory demographic calibration of S-ANT1 showed a numerically higher AUROC in Taiwanese Hokkien speakers. These findings are exploratory and hypothesis-generating, necessitating validation in larger independent cohorts before clinical implementation.
LifeSciBench is introduced, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work, with each constituent task paired with a human expert-written rubric.
Amelia Liu, Andrew Ho, Anne Marie Droste et al.· bioRxiv· 2 citations
TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations, is proposed and partial model tomography is introduced, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations.
Arooj Arif, T. Hartung, E. Botoeva et al.· 1 citation
A rapidly advancing precision-therapy pipeline-including antisense oligonucleotides to upregulate the intact allele, AAV-based gene replacement, CRISPR-mediated transcriptional activation, epigenetic modulators, and rational pathway-targeted small molecules-offers realistic prospects for disease modification.
This paper proposes the Evaluation Context Protocol (ECP), an early-stage, vendor-neutral framework intended to act as a portable evaluation contract layer for agentic systems and describes an open-source reference implementation that includes adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI.
Large Language Models show potential in their diagnostic accuracy and consequent ability to reduce clinician burden, and may provide the greatest benefit when used to optimise referral quality at source, improving both clinician and potentially LLM triage downstream.
K. Surendran, I. Aziz, Glyndwr Jenkins· Current Surgery Reports· 0 citations
Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [≥]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.
H. Bergman, V. Liu, B. Austin et al.· medRxiv· 0 citations
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.