This study is the first to propose a dual-layered engineering blueprint for AI-native applications, revealing that AI-native applications are distinguished by two core pillars: the central role of AI as the system's intelligence paradigm and their inherently probabilistic, non-deterministic nature.
Abstract
Background: The rapid advancement of large language models (LLMs) has given rise to AI-native applications, a new paradigm in software engineering that fundamentally redefines how software is designed, developed, and evolved. Despite their growing prominence, AI-native applications still lack a unified engineering definition and architectural blueprint, leaving practitioners without systematic guidance for system design, quality assurance, and technology selection. Objective: This study seeks to establish a comprehensive understanding of AI-native applications by identifying their defining characteristics, key quality attributes, and typical technology stacks, as well as by clarifying the opportunities and challenges they present. Method: We conducted a grey literature review, integrating conceptual perspectives retrieved from targeted Google and Bing searches with practical insights derived from leading open-source projects on GitHub. A structured protocol encompassing source selection, quality assessment, and thematic analysis was applied to synthesize findings across heterogeneous sources. Results: We finally identified 106 studies based on the selection criteria. The analysis reveals that AI-native applications are distinguished by two core pillars: the central role of AI as the system's intelligence paradigm and their inherently probabilistic, non-deterministic nature. Critical quality attributes include reliability, usability, performance efficiency, and AI-specific observability. In addition, a typical technology stack has begun to emerge, comprising LLM orchestration frameworks, vector databases, and AI-native observability platforms. These systems emphasize response quality, cost-effectiveness, and outcome predictability, setting them apart from conventional software systems. Conclusion: This study is the first to propose a dual-layered engineering blueprint...
A Systematic Mapping Study on the quality of AI-based software identifies six recurring challenge categories, with the most prominent being limitations in existing quality assessment models followed by issues in non-functional requirement management, quality-aware development, and quality assurance.
Maryum Hamdani, Mateen Ahmed Abbasi, Marko Jäntti et al.· 0 citations
The recent meteoric rise of LLMs (Large Language Models) and associated tools was largely unexpected and surprising to most. The rapid ascent of this technology has caught many software developers unawares, leaving them suddenly somewhat ignorant, and arguably under-skilled.
LLMs, whilst still advancing, have recently demonstrated impressive capabilities in their ability to assist software developers in their day-to-day tasks (e.g., coding new features, and locating and fixing issues). However, the use and adoption of LLMs presents many larger challenges for society as a whole; many of which are not in themselves technical concerns.
This paper examines the current and perceived impact of this technology in the context of Open Source. We identify several social, economic, environmental, political, legal, and technical concerns regarding the use of LLMs in Open Source projects.
We contribute guidance around defining an AI Policy for Open Source projects. We further offer an AI Policy Score Card to assist projects in clearly defining and declaring how they wish to work with AI or not.
Adam Retter· Balisage Series on Markup Te...· 0 citations
Artificial Intelligence (AI), particularly generative AI based on Large Language Models (LLMs), has rapidly transformed the execution of knowledge-intensive activities across multiple domains. AI-powered tools such as ChatGPT, GitHub Copilot, Microsoft Copilot, Google Gemini, Claude, and Notion AI have increasingly been adopted to automate repetitive tasks, support decision-making, accelerate software development, and improve content production. However, despite the rapid expansion of these technologies, scientific evidence regarding their effectiveness as productivity enhancers remains distributed across different research areas. This study presents a systematic literature review aimed at synthesizing current evidence on the role of AI tools in improving productivity in business, education, software engineering, and scientific research. A structured literature search was conducted across major academic databases, including Google Scholar, IEEE Xplore, ACM Digital Library, ScienceDirect, SpringerLink, and Scopus. Studies published between 2020 and 2026 were analyzed according to predefined inclusion and exclusion criteria. The reviewed literature indicates that AI tools can improve productivity by reducing task completion time, assisting knowledge creation, supporting programming activities, and enhancing information processing. Nevertheless, significant challenges remain, including inaccurate outputs, algorithmic bias, privacy concerns, ethical risks, and excessive reliance on automated systems. The findings suggest that AI achieves the greatest productivity benefits when used as a collaborative technology that augments human capabilities rather than replacing human expertise. Future research should investigate long-term productivity impacts, organizational adaptation strategies, responsible AI governance, and the integration of advanced multimodal AI systems into professional workflows.
Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
AI-generated C++ code has a distinct quality profile, showing higher rates of interface and coupling burdens, copy and allocation overheads, and a reliance on explicit loops over optimized standard APIs, which translates into tangible downstream costs, including increased review effort and a 5-8% increase in compute resource consumption.
Michael Tran, Fred Lewis, Kun Yang et al.· 1 citation
This review highlights the field's strong interdisciplinary character and reveals the current challenges that the sector faces, and outlines a systematic research program to guide further studies of the implementation, impact, and problems of GenAI in enterprises and communities.
Majdouline Attaoui, Wissal Attaoui, Anas Moukrim et al.· International journal of mul...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.