This paper introduces the notion of Hybrid Intelligence Effort (HIE), conceptualizing effort as the combined burden of model-performed reasoning and human oversight activities, and compares the explanatory power of traditional estimation metrics against interaction- and oversight-based Hybrid Intelligence dimensions.
Abstract
Software effort estimation remains a cornerstone of project planning and control, yet existing estimation models are grounded in the assumption that software development effort is dominated by human reasoning and manual construction. The rapid integration of large language models (LLMs) into development workflows fundamentally challenges this assumption by automating substantial portions of code synthesis while shifting human effort toward supervision, validation, and integration. As a result, traditional effort estimation proxies such as Story Points and size-based metrics may no longer reliably characterize development effort. This paper presents an empirical study examining how effort manifests in LLM-assisted software development. Rather than using LLMs as predictive estimation tools, we investigate how their adoption reshapes the underlying cost structure of development work. We introduce the notion of Hybrid Intelligence Effort (HIE), conceptualizing effort as the combined burden of model-performed reasoning and human oversight activities. Using a controlled experiment involving 22 developers, 110 real-world tasks, and three LLMs, we compare the explanatory power of traditional estimation metrics against interaction- and oversight-based Hybrid Intelligence dimensions. Our results show that while Story Points retain partial explanatory validity, they fail to capture dominant sources of effort in LLM-assisted workflows. In controlled experiments, HIE dimensions increase explained variance in observed effort from approximately 72–80%, while substantially reducing systematic estimation error. Human validation and corrective intervention emerge as the primary drivers of effort, outweighing artifact-level characteristics. These findings suggest that effort estimation models must move beyond human-centric and size-based assumptions to remain effective in AI-augmented software engineering.
A three-level taxonomy inspired by autonomous driving that distinguishes degrees of autonomy along a roadmap from today’s AI-assisted development workflows to fully autonomous software development in which AI systems autonomously identify demands and design, implement, verify, and maintain software without human oversight is introduced.
ACEM (Agentic Cost Estimation Model), which decomposes total agentic development cost into three additive dimensions: LLM, HITL, and infrastructure cost, is presented as a fully specified model structure and calibration methodology, with constants left symbolic pending empirical grounding.
The case suggests that generative AI is especially useful when requirements are only partially formalized, yet objective feedback from tests, benchmarks, and model quality metrics is available, and the results suggest that AI-augmented development is a relevant topic for scientific software engineering.
Robin Nunkesser· International Conference on...· 0 citations
A student survey study is presented that examines perceptions of LLM output understanding, validation effort, trust and the perceived usefulness of vibe modeling across several AI-assisted development scenarios to inform future studies for trustworthy and explainable AI-based software engineering via vibe modeling.
Shalini Chakraborty, M. Mittermaier, Judith Michael· arXiv.org· 0 citations
AI-assisted development tools enable software engineers to generate implementations at substantially higher speed and volume than in traditional workflows, yet relatively little is known about how existing guardrails evolve in response.
The growing adoption of Large Language Models (LLMs) in Software Engineering has reinforced the expectation that coding activities can be largely automated. However, this perception may represent yet another historical search for a solution capable of eliminating the inherent challenges of software development. This article discusses the transition from a code-centered paradigm to Specification-Driven Development. We argue that artificial intelligence reduces some of the effort associated with writing source code, but it does not eliminate the complexity of developing professional software systems. Instead, it shifts this complexity toward domain understanding, requirements elicitation, specification development, validation, maintenance, and software evolution. Building on this perspective, we discuss the renewed centrality of Requirements Engineering, considering its implications for productivity and software quality, as well as risks associated with automation bias, ambiguity propagation, Specification Overfitting, and the accumulation of Specification Debt. Finally, we propose the Specification Paradox: the more capable artificial intelligence systems become at automatically generating software, the greater the dependence on correct, complete, verifiable, and explainable human-produced specifications. We conclude that the future of Software Engineering will depend not only on machines'ability to generate code, but also on humans'ability to correctly specify, evaluate, and evolve what is intended to be built.
T. Sirqueira, Jessica Faciroli· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.