It is asserted that trustworthiness is a systems property, not a single algorithmic feature, and that realising it requires coordinated advances in explainability, uncertainty quantification, human–AI interaction design, failure detection, and domain-specific data governance.
The reality of deploying artificial intelligence in safety-critical systems, such as autonomous vehicles, medical diagnoses and weather forecasting, is considered, including how an AI's mathematical properties relate to its benefit and risk profile.
T. Kolda· Philosophical transactions....· 1 citation
Synthesizing across four constraint-bound engineering domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.
Omar Al-Refai, Ibrahim Shahbaz, A. Husseinat et al.· arXiv.org· 1 citation
It is argued that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms.
Timothy Merritt, Alejandro Jarabo-Peñas, Juan Bravo-Arrabal et al.· 0 citations
The surging deployment of Artificial Intelligence (AI) systems across safety-critical sectors — namely healthcare diagnostics, autonomous driving, aircraft regulation, and industrial automation — has resulted in a burgeoning need for transparent decision-making frameworks capable of justifiability and accountability. This paper describes a holistic architectural paradigm for constructing Explainable Artificial Intelligence (XAI) systems which provide sufficient reasoning transparency in mission-critical settings, where the failure of a system may have disastrous outcomes. We examine the limitations of black-box AI today and introduce a layered explainability framework that merges post-hoc explanation methods with attention processes and methodologies for synthetic formation of human personified I/O logic rules. It builds on SHAP (SHapley Additive exPlanations), LIME (Local Interpretable Model-agnostic Explanations) and counterfactual reasoning to provide fine-grained, context-aware explanations for model predictions. We also align with regulatory compliance needs, including the EU AI Act and FDA guidelines, making explainability an integral part of our design process rather than a standalone consideration. We conduct extensive experimental evaluations over medical imaging, autonomous driving, and fault detection datasets, showcasing that our architecture yields comparable predictive accuracy while substantially improving interpretability scores. In summary, this work addresses the key challenge of the disparity between AI performance and human trust by providing a solid underpinning for responsible deployment of AI in life-critical settings.
Sajjan Choudhuri, A. Agade, Rahul Reddy Gouravaram et al.· 2026 International Conferenc...· 0 citations
AI alignment is generally associated with ethics and social aspects. For high-risk AI systems, principles such as safety, stability and performance are crucial for their adoption in real-world application, to build trust and increase efficiency of a process. By safety it is meant both safety of the system itself but also safety of the process and environment. Any decision taken by a high-risk AI system should preserve safety, ensure continuous operation, real-time functioning, smooth control, and improve process efficiency. Human oversight needs to be continuously ensured, both to preserve safety of the process in case of transition from autonomous to manual mode in case of failures, but also to allow for contextual knowledge of the decisions. All these objectives are included in the AI operational alignment. High-risk AI systems designed to automatize complex processes in critical environments are continuously adapting to dynamic contexts, thus a static alignment might not be sufficient to properly assess their behaviour. The paper describes an approach to automated dynamic operational alignment in a complex high-risk process, exemplified by a case study on autonomous drilling. Possible sources of AI misalignments in this case are discussed and their potential implications.
Rodica Mihai, B. Daireaux, E. Cayeux· Scientific Reports· 0 citations
IRED-AI is an intelligent framework designed for the early detection of AI-induced risks in safety-critical systems, including healthcare, autonomous driving, aviation, industrial control, and cybersecurity. Unlike traditional detection models that optimize single metrics, IRED-AI unifies continuous monitoring, hybrid statistical–AI anomaly detection, explainable reasoning, and retrieval-guided adaptation into a cohesive, SLA-compliant architecture. A multi-view feature engineering pipeline captures contextual, signal, and semantic health indicators to construct risk embeddings that enable real-time anomaly scoring and efficient case retrieval. The framework’s adaptive alert escalation process provides operators with actionable insights at informational, cautionary, critical, and emergency levels, while integrated explainable AI modules ensure transparency and regulatory compliance. Best results across ten benchmark datasets confirm its superiority: accuracy of 95.20%, AUROC up to 0.960 (aviation), AUPRC gains of +3.37%, latency reduced to 90–150 ms (−20.83%), false alarms lowered to 7.20% (−22.22%), throughput reaching 1500 q/s, Top-1 retrieval accuracy of 92.40% (Top-5 = 97.20%), robustness of 0.920, interpretability of 0.900, and SLA compliance of 98.00%. Statistical analysis confirms these improvements as significant (p < 0.050), demonstrating that IRED-AI’s enhancements are not incidental but systematically robust. The framework also incorporates a knowledge retention loop, enabling continuous adaptation through operator feedback and periodic recalibration, ensuring resilience against data drift, adversarial manipulation, and emerging risks. By combining technical excellence with practical trustworthiness, IRED-AI provides a deployable, domain-general solution that advances safety assurance in mission-critical environments where failures carry unacceptable human and economic consequences.
Thacha Lawanna· Engineering Journal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.