Skip to content
#small language model Open access

Efficiency, Cost, and Quality: Estimating Token Consumption as a Process Metric for AI-Native Software Projects

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Efficiency, Cost, and Quality: Estimating Token Consumption as a Process Metric for AI-Native Software Projects Author: Wenchao Shi Date: September 2026 Keywords: AI-native development, token accounting, process metrics, cost of quality, statistical process control, ISO 9001, COCOMO, evidence-based management Abstract AI-native software development delegates execution to large language model agents, and its dominant variable cost is billed computation measured in tokens. Project managers nonetheless lack a defensible way to turn token counts into judgements about efficiency, cost, and quality. This paper argues that token consumption can serve as the AI-native analogue of the process input measures that classical quality management already trusts, and it specifies a protocol that makes the analogy operational. The protocol derives its unit of account from session records (input, output, reasoning, cache-read and cache-write tokens, model identity, and parent linkage); closes project attribution through a four-level rule hierarchy that ends in an explicit review queue rather than a guess; separates billed, estimated, and unknown cost registers; segments consumption by milestone with an additive conservation check; and reports a small indicator set with empirically derived thresholds. The argument for admissibility is built on the classical canon rather than on convenience: goal-derived measurement (goal-question-metric), statistical process control and the distinction between common and assignable causes, the cost-of-quality taxonomy of prevention, appraisal, and failure, the process approach and evidence-based decision making of ISO 9001, the process performance baselines of CMMI, lean waste analysis, and the effort-based cost models of software engineering (COCOMO and function points). Under these criteria, token consumption is an internal input attribute that measures effort consumed rather than value delivered, and it earns interpretive power only when paired with outcome gates, uncertainty disclosure, and explicit safeguards against gaming. A de-identified illustration from a fourteen-day embedded-systems project (approximately 790 million tokens across four milestones) shows what the indicators reveal, and which classification and pricing pitfalls corrupt them. The paper closes with the limits of the metric and the conditions under which it should not be used to judge people or projects. 摘要: AI 原生开发把执行交给大语言模型 Agent,其主要变动成本是以 token 计量的算力账单,但项目管理者至今缺少一种可辩护的方式,把 token 数量转化为对效率、成本与质量的判断。本文主张:token 消耗可以充当那些经典质量管理已经信任的"过程投入量"指标在 AI 原生场景下的对应物,并给出一套让这一类比可操作的协议。该协议以会话记录为计量单位(输入、输出、推理、缓存读、缓存写五类 token,模型身份与父会话链),用四级归类规则闭环项目归属并以显式待复核清单收尾而非猜测,严格区分账面、估算与未知三种费用口径,按里程碑分段核算并以加和守恒自检,最后报告一小组带实证阈值的指标。合法性的论证建立于经典质量管理文献而非便利:目标导向的度量(目标-问题-指标)、统计过程控制与普通原因/可归因原因的区分、预防—鉴定—失效的成本分类、ISO 9001 的过程方法与循证决策、CMMI 的过程绩效基线、精益的浪费分析,以及软件工程的基于工作量的成本模型(COCOMO 与功能点)。在这些判据下,token 消耗是一种衡量"投入了多少工作量"的内部属性,而非衡量"交付了多少价值",只有在配合结果门槛、不确定度披露与反博弈约束时才具有解释力。文中给出一个去标识化的十四天嵌入式项目实例(约 7.9 亿 token,四个里程碑),展示这些指标能揭示什么,以及哪些归类与计价陷阱会破坏它们,最后给出该指标的适用界限与不应使用它的情形。 Keywords: AI-native development; token accounting; process metrics; cost of quality; statistical process control; ISO 9001; COCOMO; evidence-based management 1. Introduction Software organisations have always estimated and controlled their work through input measures. Person-months, function points, story points, build minutes, and cloud spend are all proxies for effort, and each has survived because it is measurable, comparable within a context, and actionable. AI-native development, in which large language model (LLM) agents perform the bulk of implementation work under human direction, introduces an input measure that is unusually precise and unusually opaque at the same time. It is precise because every model interaction is metered: input tokens, output tokens, reasoning tokens, cache reads, and cache writes are counted by the serving infrastructure and recorded in agent session stores. It is opaque because no established methodology says what those counts mean for project efficiency, cost, or quality. Three practical pressures make the question urgent. First, cost: in agent-driven work the dominant variable cost is computation, and it is billed per token, so the consumption figure is simultaneously an engineering and a financial quantity. Second, capacity: token budgets bind long-horizon projects in the same way that headcount binds traditional ones, and teams that cannot explain where tokens went cannot plan. Third, and least discussed, quality: an unknown but large share of token consumption is rework, recovery, and meta-work (searching history, re-establishing context, retrying failed operations), which classical quality management has a name for and a way to interpret. The temptation is to treat tokens as a simple cost number and stop there. This paper argues for a more disciplined position: token consumption is admissible as a process metric in the sense that the quality-management tradition uses that term, provided it is defined with a goal, computed with a documented accounting protocol, interpreted through the cost-of-quality taxonomy and the logic of statistical process control, and protected against the pathologies that afflict every effort metric. The contribution is not a new statistic but a specification: what must be true of a token figure before a manager may reason from it. The paper is organised as follows. Section 2 sets out the classical foundations that supply the admissibility criteria, drawn from measurement theory for software, statistical process control, cost-of-quality accounting, the ISO 9001 process approach, CMMI process performance, lean waste analysis, Six Sigma defect logic, and software cost estimation. Section 3 specifies the token-accounting protocol: data source, attribution closure, cost registers, milestone segmentation, indicator set, and reporting discipline. Section 4 develops the central argument, mapping each element of the protocol onto a criterion from Section 2 and stating where the analogy fails. Section 5 consolidates the indicators in a definition table. Section 6 reports a de-identified illustration from a fourteen-day embedded project and the pitfalls that corrupted its first accounting. Section 7 examines threats to validity, and Section 8 concludes with the conditions under which the metric should not be used. Two boundaries are fixed at the outset. This paper does not claim that token consumption measures value, productivity in the economic sense, or individual performance. And it does not claim that the thresholds reported by any single implementation are universal: they are baselines, in the CMMI sense, valid for the process that produced them until a broader sample contradicts them. 2. Classical foundations and the criteria they impose 2.1 Measurement must be derived from a goal The first criterion comes from the goal-question-metric (GQM) tradition introduced by Basili and Weiss and formalised by van Solingen and Berghout: a metric is legitimate only when it descends from an explicit goal, through questions that operationalise the goal, to the data that answer those questions. The order matters. Metrics chosen first and justified afterwards are, in this tradition, measurement without a purpose, and they reliably drift into surrogate objectives. Applied here, the goal is not "reduce tokens". A defensible goal statement is closer to: understand, for a given project, how much execution effort each milestone consumed, whether that consumption was stable or driven by identifiable disruptions, how much of it was spent on prevention, appraisal, and failure respectively, and whether the project's cost trajectory is consistent with its remaining scope. Efficiency, cost, and quality are then three questions posed to the same data set, and each requires different comparisons: efficiency compares consumption against delivered, accepted increments; cost compares consumption against a price list, with explicit registers; quality compares the composition of consumption against the cost-of-quality categories. Fenton and Pfleeger's distinction between internal and external attributes supplies the necessary modesty. Token counts are internal attributes of the development process, observable directly from the session store. Productivity, reliability, and maintainability are external attributes that depend on context and on judgements made outside the measurement. A methodology that treats an internal attribute as if it were an external one produces the specific failure this paper tries to prevent. 2.2 Statistical process control: variation before judgement The second criterion comes from Shewhart and from the practice Deming built on it. A process metric is interpretable only when the process is in a state of statistical control, that is, when variation arises from common causes inherent to the process rather than from sporadic assignable causes. Shewhart's control chart was designed to make that distinction visible, and Deming's insistence that most variation is systemic (the famous attribution of the large majority of problems to the system rather than to the worker) follows from it. For token accounting the consequence is concrete. A single day of unusually high consumption is not a performance signal until the analyst asks whether the process has a stable baseline and whether the excursion has an identifiable cause: a failed approach that had to be retried, a recovery session after a lost context, a large-scale refactor, or a data-classification error. The protocol in Section 3 therefore reports daily consumption, the share of the peak day, and the share of meta-work sessions, and it treats these as candidate assignable causes rather than as verdicts. Deming's related warnings, against management by numbers alone and against driving fear into the system, also bound the use of the metric: consumption data used to rank individuals converts a diagno

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.