Skip to content

Author

Ethan Williams

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2020

Cloud-Native Data Pipelines for Enterprise Analytics

Cloud native data pipelines have become an enabling ingredient of the modern enterprise analytics to fulfill the ever-increasing demand of a scale-loving, resilient, and real-time processing of a wide range of data sources. Organizations currently produce large amounts of structured, semi-structured, and unstructured data in transactional systems, Internet of Things (IoT) platforms, digital channels, and data sources that are external (Bank of America 2017). Old monolithic data integration architectures are designed to provide batch-oriented processing and static infrastructure capabilities have challenges satisfying low latency, scale on demand, and 24/7 requirements. Reactively, cloud-native paradigms, including the foundations of microservices, container orchestration, event-driven architectures, and managed cloud services have caused a rethinking of the data pipeline design, deployment and operation. This article provides an in-depth analysis of cloud-native pipeline data to enterprise analytics along with their main architectural concepts and processing models as well as operational aspects that fall within the professional scope of IEEE publications. The research paper summarizes the literature and business methodologies to present a reference model which brings together data ingestion, stream processing, batch processing, storage, governance, and analytics consumption layers. Special concern is opened to the contributions of containerization, orchestration platforms, and serverless computing towards facilitation of elasticity and fault tolerance. The paper also examines design patterns like Lambda architecture and Kappa architecture, data mesh theory and metadata-based orchestration, with an emphasis on its application to the large enterprise environment. An organized approach to the design and deployment of cloud-native data pipelines with the inclusion of data quality management, security controls, observability, and cost optimization is suggested. Throughput, latency and scalability modeling mathematical formulations are proposed in order to facilitate capacity planning and performance measurement. Representative enterprise workloads as shown through experiment results exhibit evident increases in data processing latency, pipeline reliability and operational efficiency over traditional architectures. These findings are placed in context to the discussion of the broader transformation efforts at enterprises, whereas the conclusion provides recommendations on future research opportunities, such as autonomous pipeline optimization and AI-based orchestration.

Ethan Williams · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.