Skip to content
Open access

Scalable Software Engineering Architecture for AI-Enabled ETL Pipelines Using Event-Driven Microservices

Jul 2026 · ICCK Journal of Software Engineering · Vol 2, pp. 169-184 · 0 citations

TL;DR

An event-driven microservices approach to orchestrate and deploy AI-based ETL pipelines in a Kubernetes-managed environment that includes the following components: asynchronous orchestration using Apache Kafka, hybrid anomaly detection, adaptive schema inference, and predictive load balancing.

Abstract

As data ecosystems become more diverse and time-critical, traditional monolithic ETL pipelines face challenges to meet the demands of modern data engineering workloads in terms of scalability, adaptability, and operational resilience. In this paper, we introduce an event-driven microservices approach to orchestrate and deploy AI-based ETL (ETL = Extraction, Transformation, and Loading) pipelines in a Kubernetes-managed environment that includes the following components: asynchronous orchestration using Apache Kafka, hybrid anomaly detection, adaptive schema inference, and predictive load balancing. The proposed architecture breaks the ETL processing into loosely coupled services, which can be deployed and scaled independently, and incorporates data quality intelligence into the transformation layer. Two explicit baselines are used for experimental evaluation: (1) a traditional monolithic batch-processing ETL pipeline with sequential execution of the stages performed without horizontal scaling, and (2) an event-driven microservices pipeline, which uses Apache Kafka orchestration but no AI-enabled optimization modules. The framework is able to achieve 39.7% improvement in throughput (4,820 vs. 3,450 records/sec) and 96.3% anomaly detection accuracy, while reducing average E2E latency by 14.6% (245 vs. 287 ms). The framework's throughput is 288.7% higher than the baseline of ETL monolithic, and the average latency drops 72.5% to 245 ms compared to 892 ms for the baseline. An architectural assessment also shows that the system is more modular, loosely coupled, more fault isolated, and more flexible to deploy; all of which are important software quality attributes. The results indicate the proposed architecture as a potentially reusable reference to create scalable and intelligent ETL systems in enterprise data processing environments and emphasize the necessity of generalization in multi-domain validation.

Read PDF

Similar papers

Open access Sep 2026

Serverless Data Engineering: Innovations in Python-Driven ETL Automation on AWS

Traditional cluster-based ETL architectures impose a structural tax on data engineering organisations: fixed compute resources provisioned for peak demand, scheduled batch cycles that introduce latency regardless of downstream urgency, and operational overhead that redirects engineering capacity from pipeline design to...

Rambabu Bolineni · 0 citations
Open access 2024

AI-Assisted Data Pipeline Orchestration for Scalable Analytics

Modern enterprises face increasing demands for scalable and efficient data processing due to rapid data growth. Traditional data pipeline orchestration methods, which rely on static configurations and manual intervention, often lead to inefficiencies in resource use, latency, and fault tolerance. This paper proposes an...

J. Weizenbaum, S. Papert · 1 citation
Open access 2019

Continuous Data Transformation in Event-Driven Microservices

In modern distributed systems, event-driven microservices have emerged as a robust architectural paradigm, offering scalability, resilience, and decoupled communication. However, these systems often require real-time or near-real-time transformation of data across services and domains. Continuous data transformation—th...

Thabo Nkosi · 0 citations
Conference Jul 2026

Intelligent Data Engineering Pipelines for Enterprise Applications: Architecture Challenges and Optimization Strategies

The rapid growth of enterprise applications has led to a substantial increase in the volume, velocity, and variety of data, necessitating intelligent data engineering pipelines for efficient processing and analytics. These pipelines play a critical role in enabling scalable data integration, transformation, and deliver...

D. Bansal, Dinesh Kumar Garg · 0 citations
Open access 2020

Cloud-Native Data Pipelines for Enterprise Analytics

Cloud native data pipelines have become an enabling ingredient of the modern enterprise analytics to fulfill the ever-increasing demand of a scale-loving, resilient, and real-time processing of a wide range of data sources. Organizations currently produce large amounts of structured, semi-structured, and unstructured d...

Ethan Williams · 0 citations
Conference Jul 2026

Agentic Data Services: A Control-Plane Architecture for Adaptive Data Workflows

Modern data platforms rely on pipeline-oriented architectures that are rigid, hard to adapt, and lack native auditability. We present Agentic Data Services, a control-planedriven architecture for Big Data as a Service (BDaaS) that models workflows as adaptive, policy-aware service entities rather than static directed a...

Alexander Chernov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.