Aug 2026· Kaohsiung Journal of Medical Sciences· 0 citations· 10 references
Medicine
TL;DR
It is suggested that primary spoken language may be associated with ANT1 performance in MHE assessments and Integrating ANT1 with serum IL-6 showed numerically improved discrimination in Mandarin speakers, whereas exploratory demographic calibration of S-ANT1 showed a numerically higher AUROC in Taiwanese Hokkien speakers.
Abstract
ABSTRACT Minimal hepatic encephalopathy (MHE) involves subtle cognitive dysfunction and systemic inflammation and is associated with an increased risk of overt hepatic encephalopathy. The Animal Naming Test (ANT1) is a rapid semantic fluency tool for MHE assessment, but its performance across different primary spoken languages remains unclear. We conducted a prospective proof‐of‐concept study to evaluate the diagnostic performance of ANT1 and serum interleukin‐6 (IL‐6) in Mandarin‐ and Taiwanese Hokkien‐speaking cirrhotic patients. A total of 65 cirrhotic patients and 34 healthy controls were enrolled. Patients completed ANT1, simplified ANT1 (S‐ANT1), and standard psychometric assessments. MHE was defined by abnormal PHES and/or visually assessed EEG slowing. Diagnostic discrimination was evaluated using AUROC analyses stratified by primary spoken language. A post hoc exploratory Taiwanese‐calibrated S‐ANT1 was also assessed in Taiwanese Hokkien‐speaking patients. Sixteen cirrhotic patients (24.6%) were diagnosed with MHE. Patients with MHE had lower ANT1 scores and higher serum IL‐6 levels than those without MHE. In Mandarin‐speaking patients (n = 44), ANT1 demonstrated an AUROC of 0.760, while serum IL‐6 showed an AUROC of 0.841. The composite model combining ANT1 and serum IL‐6 showed a numerically higher AUROC of 0.895, but this improvement was not statistically significant in pairwise DeLong comparisons. In Taiwanese Hokkien‐speaking patients (n = 21), standard ANT1 showed a lower AUROC of 0.679. The exploratory Taiwanese‐calibrated S‐ANT1 showed a numerically higher AUROC of 0.776; however, this post hoc finding was based on a small subgroup with 7 MHE events and requires external validation. This study suggests that primary spoken language may be associated with ANT1 performance in MHE assessments. Integrating ANT1 with serum IL‐6 showed numerically improved discrimination in Mandarin speakers, whereas exploratory demographic calibration of S‐ANT1 showed a numerically higher AUROC in Taiwanese Hokkien speakers. These findings are exploratory and hypothesis‐generating, necessitating validation in larger independent cohorts before clinical implementation.
FLARE is proposed, a novel framework that endows VLAs with robust error recovery capabilities through a ``Retry" and ``Reset" Paradigm, and significantly improves task success and robustness.
Ganlong Zhao, Zijia Tang, Xingping Chen et al.· 3 citations
LifeSciBench is introduced, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work, with each constituent task paired with a human expert-written rubric.
Amelia Liu, Andrew Ho, Anne Marie Droste et al.· bioRxiv· 2 citations
Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.
Yu-Fan Wu, Yinghui He, Zhengyi Hu et al.· 1 citation
TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations, is proposed and partial model tomography is introduced, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations.
Arooj Arif, T. Hartung, E. Botoeva et al.· 1 citation
The complex multi-energy coupling characteristics inherent to integrated energy system (IES) present unprecedented challenges for the implementation of low-carbon scheduling. Existing optimization methods often exhibit limitations in system scalability, algorithm adaptivity, and carbon reduction efficacy for complex IES. This paper proposes a Large Language Model (LLM)-Embedded Multi-Agent Reinforcement Learning (LEMARL) to address the aforementioned issues. The proposed method integrates the global perception capability of LLMs with the dynamic optimization capability of MARL. Specifically, the LLM-Embedded module generates high-quality reward functions and policy frameworks from a global perspective, while the MARL module leverages these LLM-generated strategies for distributed interactive iterations—greatly enhancing computation efficiency and scalability. Simulation results demonstrate that LEMARL reduces carbon emissions by 7.76% and simultaneously decreases operating costs by 4.49% in a small-scale IES. Furthermore, LEMARL also exhibits superior applicability and scalability in large-scale IES of the IEEE 141-bus power grid integrated with 51-node thermal system.
Chen Xia, Tong Gou, Yinliang Xu et al.· IEEE Transactions on Smart G...· 1 citation
General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of executable file, search, and code environments, while \emph{Agentic Coordination Scaling} trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a \emph{Heavy-Duty Solver} for ambitious, long-running tasks.
Apodex Team B. An, B. Li, B. Wang et al.· 1 citation
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.