Skip to content
Review

On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems

Jul 2026 · arXiv.org · Vol abs/2607.23365 · 1 citation
Computer Science

TL;DR

AITD-MAP is introduced, an integrated framework that connects the AITD taxonomy, quality and risk impacts, and mitigation strategies into a unified structure for risk-aware AI engineering, and aims to assist AI software engineers in making AI safety and security technical debts visible, understanding their root causes, and mitigating their presence.

Abstract

Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education. While these systems offer powerful data-driven and adaptive capabilities, their complexity, rapid evolution, and dependence on dynamic data pipelines introduce new forms of engineering liability collectively referred to as AI Technical Debts (AITDs). AITDs arise from root causes spanning data governance, model implementation, algorithm design, architectural decisions, operational processes, documentation practices, and testing adequacy. Unlike conventional technical debt, many AITDs are latent and propagate across tightly coupled AI pipelines, leading to maintenance challenges, reliability degradation, and heightened safety or security risks. Guided by the principles of AI Trust, Risk, and Security Management (AI TRiSM), this study reinterprets technical debt through the interconnected dimensions of trustworthiness, focusing on AI safety and security technical debts. We conduct a systematic review of 60 primary studies and identify 31 distinct types of AITD, which are organized into a root-cause-oriented taxonomy comprising seven classes. The analysis examines how these debts map to 18 trust-related concerns, including 6 safety hazards and 12 security vulnerabilities. To support mitigation, the review synthesizes 34 actionable guidelines (8 safety and 26 security) targeting the prevention, detection, and reduction of AITDs across the AI lifecycle. Building on these findings, we introduce AITD-MAP, an integrated framework that connects the AITD taxonomy, quality and risk impacts, and mitigation strategies into a unified structure for risk-aware AI engineering. The framework aims to assist AI software engineers in making AI safety and security technical debts visible, understanding their root causes, and mitigating their presence.

View source

Similar papers

Sep 2026

Governing data risks in the age of AI

Artificial intelligence (AI) and in particular generative AI (GenAI) has accelerated data risk in financial services by changing how data is accessed, transformed, and used to drive decisions that regulators closely scrutinise. The pace of AI adoption has outrun many organisations’ data control foundations: AI tools frequently require broad access to data systems, deepen reliance on third-party vendors, and create new pathways through which sensitive data can leak, via prompts, model outputs, automated retrieval processes, and AI-driven actions. At the same time, regulatory expectations for data accuracy, completeness, timeliness, and traceability remain uncompromising, particularly for high-stakes use cases such as capital and liquidity management, regulatory and financial reporting, and financial crime detection. This paper proposes a practical, audit-ready approach to governing data risks in the AI era. The central idea is to add a second lens to traditional data classification, one focused on business consequence rather than confidentiality alone. Specifically, it introduces the concept of critical data elements (CDEs): data elements whose inaccuracy, unavailability, or misuse can produce material regulatory, financial, or customer-facing impact. Pairing CDE designation with conventional confidentiality classifications creates a dual-axis model that directs the strongest controls to the highest-consequence data, even when that data may not appear sensitive on the surface. Drawing on established regulatory frameworks including BCBS 239 (risk data aggregation and reporting), SR 26-2 (model risk management, the interagency guidance issued jointly by the Federal Reserve, Office of the Comptroller of the Currency (OCC), and Federal Deposit Insurance Corporation (FDIC) in April 2026, superseding SR 11-7), and US interagency third-party risk guidance, the paper explains why AI amplifies data risk across four dimensions (privacy, security, integrity, and accountability), and presents an eight-domain data governance framework with concrete audit evidence examples. The goal is to equip compliance, risk, and audit leaders with a repeatable structure for demonstrating that AI innovation rests on controlled, auditable data foundations. This article is also included in The Business & Management Collection which can be accessed at http://hstalks.com.business/.

Xin-Fa Tu · 0 citations
#explainable ai Open access Sep 2026

Enterprise AI Architecture Governance and Risk Framework

Artificial Intelligence (AI) is transforming modern enterprises by enabling intelligent automation, predictive analytics, generative AI, autonomous decision-making, digital assistants, and real-time optimization across business processes. As organizations increasingly deploy Large Language Models (LLMs), Agentic AI, Retrieval-Augmented Generation (RAG), intelligent copilots, machine learning, and AI-driven business applications, governance has become one of the most critical success factors for sustainable enterprise AI adoption. While AI creates unprecedented opportunities for innovation, productivity, customer engagement, and operational excellence, it simultaneously introduces significant risks including model bias, hallucinations, explainability limitations, privacy concerns, cybersecurity threats, regulatory compliance challenges, intellectual property exposure, model drift, ethical considerations, and uncontrolled autonomous decision-making. Traditional enterprise governance models are insufficient because they were designed for conventional software systems rather than continuously learning intelligent systems operating across hybrid cloud environments. Organizations therefore require an Enterprise AI Architecture Governance and Risk Framework that integrates Enterprise Architecture, AI lifecycle management, cybersecurity, Business Knowledge, data governance, compliance, risk management, operational monitoring, and Responsible AI into a unified governance capability. This paper proposes a comprehensive governance framework built upon SAP Enterprise Architecture Framework (SAP EAF), TOGAF®, SAP Business Technology Platform (SAP BTP), SAP AI Foundation, SAP AI Core, SAP AI Launchpad, SAP Joule, SAP Business Data Cloud, SAP Datasphere, SAP HANA Cloud, SAP Integration Suite, SAP Cloud Identity Services, SAP Cloud ALM, SAP Build, SAP LeanIX, and SAP Signavio. The framework establishes governance across strategy, architecture, data, AI models, infrastructure, security, operations, compliance, and enterprise risk while embedding continuous monitoring, explainability, lifecycle governance, and policy enforcement. The proposed architecture enables organizations to build secure, scalable, transparent, explainable, compliant, resilient, and trustworthy AI ecosystems that accelerate innovation while minimizing enterprise risk and supporting long-term digital transformation.

Sanjeeve Kumar Gajadi · 0 citations
Preprint Aug 2026

Insurance as AI Risk Infrastructure: A Generative-Agent Simulation of AI Adoption

The rapid evolution of artificial intelligence (AI) tools has demonstrated immense potential to enhance societal well-being and operational efficiency. However, the inherent unreliability and uncertain operational consequences of modern AI systems, typified by large language models (LLMs), have created a significant barrier to enterprise adoption. Many enterprises remain hesitant to integrate these tools deeply into their workflows due to concerns about unpredictable losses and liability exposure. While existing technical safeguards primarily seek to reduce the likelihood or severity of AI-enabled workflow failures, they do not by themselves provide ex post financial protection when residual pecuniary tail losses materialize. In this paper, we introduce a socio-economic framework that complements these safeguards by transferring and absorbing the residual financial consequences of AI adoption through insurance. To evaluate this framework, we develop an LLM-driven agent-based social simulation (LABSS) system. We assess the behavioral validity of the simulation using established economic and sociological theories. Our analysis demonstrates that the proposed insurance framework reduces firm-level financial exposure, thereby accelerating the aggregate adoption of AI tools and improving firm solvency and aggregate capital.

Yi-Xuan Yuan, Dedai Wei, Chudong Qian et al. · 0 citations
Review Open access Aug 2026

AI-Assisted Test Execution as an Augmentation Layer in Enterprise Quality Engineering

Artificial intelligence has introduced new capabilities for enterprise quality engineering, yet the absence of structured governance frameworks limits its practical adoption in regulated software environments. This paper presents the Probabilistic Augmentation and Governance Model (PAGM), a three-tier framework that formally allocates responsibility between AI agents and human testers across autonomous execution, confidence-gated escalation, and human-led verification. PAGM draws on a structured review of literature spanning AI quality assurance, autonomous system testing, UI-resilient automation, explainable AI, and defect prediction to identify governance gaps in existing approaches. The model addresses eight documented AI testing challenges, including interpretability, absence of formal specifications, oracle determination, and dynamic operational environments. Unlike prior approaches that treat AI test execution as a performance optimization problem, PAGM positions governance, traceability, and bounded autonomy as first-class design requirements. Its application is demonstrated through a regulated property insurance release pipeline in which PAGM enables a full quarterly regression cycle within a five-day service level agreement while preserving mandatory human accountability for premium calculation and claims adjudication workflows. Quantitative evidence from the literature supports the model's constituent mechanisms, including a 95% UI change handling rate for contextual recognition methods compared to 40 to 80 percent achieved by conventional tools and AI-based classifiers outperforming classical approaches by 25 to 50 percent across software quality metrics. PAGM provides a governance-ready foundation for organizations seeking to deploy AI-assisted testing at enterprise scale within regulated or auditable environments.

Rejenish Kiran · 0 citations
Open access Aug 2026

AI-Augmented DevSecOps for Protecting U.S. Enterprise Software Supply Chains and Critical Digital Services

Modern enterprise applications rely on extensive third-party code, automated build systems, cloud-native infrastructure and rapidly changing vulnerability intelligence. Security controls are then spread out across the development, the software supply-chain assurance and production operations making it hard to relate some process deficiency with the subsequent impact seen in operations. This paper proposes an AI-driven DevSecOps approach focusing on governance with connections between Secure Software Development, Software Provenance, Runtime Observability, AI for analysis, Deterministic Policy Enforcement, and Responsible Human Decision Making. The framework is developed in a design-science methodology, and structured as the following layers: mission and regulatory context; AI-augmented DevSecOps pipeline; cloud-native runtime; unified observability and threat intelligence; AI intelligence and decision support; and policy governance with human oversight. It has a core principle that AI can correlate evidence, assess risk, and inform action but cannot override the requirements of policy, or authorize high impact operational decisions. A healthcare software-supply-chain scenario is the basis for an example of - a controlled rollback enabled by software bills of materials, provenance records, runtime telemetry, vulnerability intelligence, AI reasoning, policy checks, and human approval. It provides an end-to-end conceptual architecture, a decision model driven by policy, and a standards-based foundation for implementation. Since the evaluation is scenario based, the framework should be viewed as a design artifact and it is necessary to prototype and empirically validate it across multiple sectors.

Mir Fawad, Mir Jawad Yaqoob, Khawar Muhammad Saad · 0 citations
Preprint Jul 2026

Studying, Identifying, and Fixing Hidden Technical Debt in AI-Intensive Cyber-Physical Systems

Artificial Intelligence (AI) components are increasingly pervasive in several software systems, including Cyber-Physical Systems (CPSs). AI-CPS are used in several domains, including autonomous vehicles, industry, home automation, robotics, and healthcare. Being composed of hardware, AI components, and conventional modules, AI-CPS can exhibit technical debt (TD) that is peculiar and potentially more challenging than that of conventional systems. This thesis aims to characterize AI-CPS TD and propose approaches for its identification and repair. In a first phase, we characterize AI-CPS TD by analyzing AI ecosystems and AI-CPS repositories, as well as interviewing developers. Based on the acquired knowledge, we define approaches to identify and mitigate such TD. Finally, we plan to develop and validate an automated tool that supports agentic AI solutions to monitor, govern, and repay AI-CPS TD.

Beena · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.