Architectural and Governance Approaches to Controlling Large Language Model Costs in Healthcare Information Systems
Abstract
Large language models are entering healthcare information systems through documentation, retrieval, summarization, multimodal interpretation, and administrative workflows. Their operating costs depend on more than unit pricing because repeated context transfers, unnecessary model calls, retries, long outputs, and reuse failures accumulate across routine work. This article examines how architectural controls and enterprise governance can contain that expenditure without weakening workflow value. Fourteen sources covering healthcare multimodality, AI economics, payer operations, model routing, prompt compression, semantic caching, inference serving, and governance were compared and synthesized. The resulting framework separates cost control into task qualification, context management, model selection, bounded generation, reusable memory, and financial accountability. It also connects expected returns agreed before deployment with operational monitoring and the following planning cycle. Payer-specific application is examined through claims payment integrity and appeals and grievances, where incremental claims value, labor capacity, compliance performance, and model consumption must be evaluated against the existing human process. The proposed approach helps healthcare organizations assign expensive reasoning only to tasks that require it, preserve validated outputs for later use, and evaluate consumption against clinical, operational, and financial objectives throughout the year.