Skip to content
#small language model Review Open access

Automating cost-effectiveness models with agentic artificial intelligence: Case study and implications for value assessment.

Sep 2026 · Journal of Managed Care & Specialty Pharmacy · Vol 32 9, pp. 1090-1100 · 0 citations · 24 references
Medicine

TL;DR

Findings support a hybrid paradigm in which AI augments, but does not replace, health economists in value assessment and formulary decision support within managed care settings.

Abstract

Background

Health economic modeling is conceptually sophisticated but operationally repetitive and resource intensive. Recent advances in large language models suggest potential for automating components of cost-effectiveness model development.

Objective

To evaluate whether an agentic artificial intelligence (AI) system can reliably automate cost-effectiveness model development in the context of targeted therapies for anaplastic lymphoma kinase-positive (ALK+) non-small cell lung cancer (NSCLC).

Methods

We developed the Agentic Health Economic Modeling Platform (A-HEMP) to construct a cost-effectiveness model for ALK+ NSCLC therapies without access to existing models in that clinical context. Modeling decisions and extracted parameters were compared with a previously published manual cost-effectiveness analysis. A-HEMP automated PICO-based scoping, modeling approach recommendation, systematic literature review, and structured parameter extraction. Performance was evaluated across 3 domains: model structure concordance, evidence identification concordance, and parameter value alignment. Deterministic cost-effectiveness outputs were calculated externally for benchmarking.

Results

A-HEMP identified a modeling framework aligned with the published cost-effective analysis and retrieved all primary clinical trials used for clinical efficacy inputs. Concordance in model structure and assumptions was observed in 27% of modeling dimensions, with 36% partially concordant and 36% divergent. Divergences were most prominent in survival extrapolation and intracranial progression handling. Evidence identification concordance was high for primary clinical trials (100%) but moderate for cost inputs, with 25% of evidence domains fully concordant and 38% partially concordant. Parameter value alignment was high for clinical efficacy inputs and progression-free health state utilities (<2% deviation), whereas greater variability was observed for sicker health states (12% deviation) and downstream disease management costs (25%-55% deviation).

Conclusions

Agentic AI can reliably automate upstream components of cost-effectiveness model development. However, nuanced modeling decisions with less standardized methodological guidance remain areas requiring expert oversight. These findings support a hybrid paradigm in which AI augments, but does not replace, health economists in value assessment and formulary decision support within managed care settings.

Read PDF

Similar papers

#small language model Open access Aug 2026

LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences

LifeSciBench is introduced, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work, with each constituent task paired with a human expert-written rubric.

Amelia Liu, Andrew Ho, Anne Marie Droste et al. · 2 citations
#small language model Preprint Aug 2026

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.

Yu-Fan Wu, Yinghui He, Zhengyi Hu et al. · 1 citation
#artificial intelligence Preprint Aug 2026

TestifAI: Tomography-Based Testing for Deep Learning Systems

TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations, is proposed and partial model tomography is introduced, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations.

Arooj Arif, T. Hartung, E. Botoeva et al. · 1 citation

Low Carbon Scheduling of Integrated Energy System Based on Large Language Model-Embedded Multi-Agent Reinforcement Learning

The complex multi-energy coupling characteristics inherent to integrated energy system (IES) present unprecedented challenges for the implementation of low-carbon scheduling. Existing optimization methods often exhibit limitations in system scalability, algorithm adaptivity, and carbon reduction efficacy for complex IES. This paper proposes a Large Language Model (LLM)-Embedded Multi-Agent Reinforcement Learning (LEMARL) to address the aforementioned issues. The proposed method integrates the global perception capability of LLMs with the dynamic optimization capability of MARL. Specifically, the LLM-Embedded module generates high-quality reward functions and policy frameworks from a global perspective, while the MARL module leverages these LLM-generated strategies for distributed interactive iterations—greatly enhancing computation efficiency and scalability. Simulation results demonstrate that LEMARL reduces carbon emissions by 7.76% and simultaneously decreases operating costs by 4.49% in a small-scale IES. Furthermore, LEMARL also exhibits superior applicability and scalability in large-scale IES of the IEEE 141-bus power grid integrated with 51-node thermal system.

Chen Xia, Tong Gou, Yinliang Xu et al. · 1 citation
#small language model Preprint Aug 2026

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of executable file, search, and code environments, while \emph{Agentic Coordination Scaling} trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a \emph{Heavy-Duty Solver} for ambitious, long-running tasks.

Apodex Team B. An, B. Li, B. Wang et al. · 1 citation

Related blog posts