Agentic Data Engineering Framework for Autonomous, Real-Time, and Trustworthy AI Systems
Abstract
The fast expansion of real-time data has revealed the very important weaknesses of traditional data engineering pipelines, especially in the aspects of flexibility, scalability, and resilience. This paper is a proposal for an Agentic Data Engineering Framework that is meant to facilitate autonomous, real-time, and trust-conscious data processing. This framework integrates a multi-agent design that is composed of decision, processing, and trust components and allows pipelines to be dynamically reconfigured based on the condition of systems and characteristics of data. The methodology builds data engineering as a multi-objective optimization and involves latency, accuracy, computational cost, and trust score as a combination of learning. A novel algorithm, Agentic Pipeline Optimization with Trust Feedback, is developed to guide adaptive decision-making, based on reinforcement learning and trust-sensitive rewards modeling. Heterogeneous datasets under controlled conditions are used to perform experimental evaluation. The proposed framework has an accuracy of 94.6 %, latency of 165 ms, throughput of 1980 records/s, and a trust score of 0.91, which is much better than the baseline models. These improvements correspond to approximately 5–10% higher accuracy, 20–30% lower latency, and 20% increased throughput compared to existing approaches. The findings indicate that the framework works well in dynamic data settings as the performance is stable. The results suggest that agentic intelligence is most likely to combine with trust-sensitive optimization to offer a viable route to next-generation data engineering systems, specifically when it comes to Internet of Things (IoT), cloud, and intelligent enterprise applications. Received: 15 April 2026 | Revised: 17 June 2026 | Accepted: 29 July 2026 Conflicts of Interest The authors declare that they have no conflicts of interest to this work. Data Availability Statement The data that support the findings of this study are openly available on the UCI Machine Learning Repository at https://archive.ics.uci.edu/dataset/352/online+retail and on the LogHub Repository (HDFS Log Dataset) at https://github.com/logpai/loghub/tree/master/HDFS. Processed datasets, experimental configurations, and supplementary implementation details are available from the corresponding author upon reasonable request. Author Contribution Statement Parikshit Premchand Sahagal: Conceptualization, Methodology, Validation, Formal analysis, Writing – original draft, Writing – review & editing, Supervision, Project administration. Swetha Talluri: Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Writing – review & editing, Visualization. Pradeep Kumar Sambamurthy: Methodology, Software, Validation, Formal analysis, Investigation, Resources, Data curation, Writing – review & editing, Visualization. Hastimal Jangid: Conceptualization, Resources, Writing – original draft, Writing – review & editing, Supervision, Project administration. Abrar Ahmed Syed: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Resources, Data curation, Writing – original draft, Writing – review & editing, Visualization, Supervision, Project administration.