Skip to content
Open access

A Hybrid Data Engineering and Generative AI Architecture for Intelligent Data Governance, Metadata Management, and Automated Data Quality Assessment on AWS

2024 · International Journal of Artificial Intelligence, Data Science and Machine Learning · Vol 5, pp. 230-240 · 0 citations

TL;DR

Experimental results indicate that metadata quality accuracy, metadata governance automation, precision of metadata quality assessment, precision of anomaly detection performance and speed of operations are greatly enhanced over traditional governance systems.

Abstract

The increasing complexity of enterprise data ecosystems has thrown new challenges at the problem of data governance, metadata management and data quality assurance. Cloud-based platforms are becoming more and more important for organizations to store, process and analyze massive amounts of structured, semi-structured and unstructured data from business applications, Internet of Things (IoT) devices, customer interactions, social media, and transactional systems. Cloud technologies offer scalable infrastructure for data management, but traditional governance practices can find it challenging to ensure high-quality data, enforce compliance policies and maintain consistency of metadata in distributed environments. The adoption of data-driven decision-making has created a critical need for more intelligent, automated and scalable governance mechanisms as enterprises go through this transition.With the transition to data-driven decision-making processes, the need for more intelligent, automated and scalable governance mechanisms has become critical. With the recent development of Generative Artificial Intelligence (GenAI), the ways in which traditional data engineering practices can be improved by augmenting them with automated metadata generation, data cataloging, data quality checks, anomaly detection, and enforcement of governance policies have expanded. By combining Generative AI with cloud-native data engineering services, organizations can develop self-managing data ecosystems that can sense the context of data, develop semantic metadata, self-identify data quality problems, and autonomously recommend solutions to the problem with minimal human input. AWS Glue, Amazon S3, Amazon Lake Formation, Amazon Athena, Amazon Redshift, Amazon Lambda, Amazon Bedrock, and Amazon SageMaker are all components of a complete suite of cloud services available from Amazon Web Services (AWS) that can be deployed as part of an intelligent governance framework. We propose a Hybrid Data Engineering and Generative AI Architecture for Intelligent Data Governance, Metadata Management and Automated Data Quality Assessment on AWS. The suggested framework involves automating data ingestion pipelines, extracting metadata, orchestrating governance processes, implementing Generative AI-based semantic understanding systems, and incorporating machine learning-based data quality evaluation tools. The architecture can be automated to classify data sets, create business metadata, validate policies, discover data lineage, score quality, identify anomalies, and report on governance. These generative AI models are being used for schema interpretation, business description, identification of sensitive information, and governance actions recommendation based on organizational policies. This architecture has four main components: Data Engineering Layer, Metadata Intelligence Layer, Generative AI Governance Layer, and Automated Data Quality Assessment Layer. Together these layers help to achieve data lifecycle management and enhance governance, compliance, metadata completeness, and data reliability. Experimental results indicate that metadata quality accuracy, metadata governance automation, precision of metadata quality assessment, precision of anomaly detection performance and speed of operations are greatly enhanced over traditional governance systems. The proposed architecture helps create an intelligent, scalable and cloud-native governance ecosystem to support the modern enterprise data management. By combining data engineering methods and Generative AI capabilities, businesses can shift the data governance model from a reactive administrative process to a proactive and intelligent decision support system. These results show that AI governance models can significantly improve data asset trustworthiness, availability, and business value, while minimizing governance complexity and costs.

Read PDF

Similar papers

Open access 2025

Large Language Model-Assisted Metadata Engineering for Enterprise Data Platforms

Enterprise metadata is essential for data discovery, governance, integration, and analytics. Traditional metadata engineering relies on manual, rule-based approaches that struggle with dynamic, heterogeneous enterprise data across cloud, IoT, ERP, CRM, and data lake environments. This study proposes a Large Language Model-Assisted Metadata Engineering Framework (LLM-MEF) that automates metadata extraction, semantic enrichment, schema recommendation, lineage discovery, and governance validation using transformer-based LLMs, Retrieval-Augmented Generation (RAG), vector databases, and knowledge graphs. The framework improves metadata quality, semantic consistency, discoverability, governance compliance, and operational efficiency while reducing manual effort. Explainable AI and continuous feedback learning further enhance transparency and adaptive improvement. The proposed LLM-MEF provides a scalable and intelligent solution for enterprise metadata management, supporting modern data governance, AI, business intelligence, regulatory compliance, and digital transformation.

David Wheeler, Michael Gordon · 0 citations
Open access 2025

Autonomous Data Fabric Architectures for Enterprise-Wide Intelligent Computing

Modern enterprises generate massive volumes of data from cloud platforms, IoT devices, enterprise applications, social media, and AI systems, creating challenges in data integration, governance, scalability, security, and real-time analytics. Traditional data management approaches often struggle to handle these complex and distributed environments. This paper proposes an Autonomous Data Fabric (ADF) architecture that combines AI/ML, metadata-driven automation, knowledge graphs, intelligent orchestration, and policy-based governance to enable seamless, self-managing enterprise data ecosystems. The framework supports automated data discovery, semantic integration, adaptive workflows, continuous monitoring, and intelligent resource optimization while ensuring data quality, security, and compliance. Experimental results demonstrate that the proposed ADF significantly improves data integration efficiency, governance, analytics performance, operational cost, and decision-making compared to conventional systems. Its scalable and self-adaptive design supports hybrid cloud, multi-cloud, edge, and on-premises environments, making it a robust solution for enterprise digital transformation and next-generation intelligent data management.

Narendra Karmarkar · 0 citations
Conference Aug 2026

Beyond ETL: A Lean AI Data Warehouse Architecture Using Model Context Protocol with Integrated Data Governance

Modern enterprises face increasing challenges in managing data infrastructure for AI-driven analytics while maintaining governance, compliance, and cost efficiency. This paper proposes a Lean AI Data Warehouse (LAIDW) architecture that integrates the Model Context Protocol (MCP) as a unifying interface layer between AI agents and heterogeneous data sources. The proposed framework reduces redundant data movement by enabling AI models to query data sources directly through standardized MCP connectors, reducing intermediate storage overhead and improving data freshness. A complementary Data Governance Framework (DGF) enforces data quality, lineage tracking, access control, and auditability across the MCP-connected ecosystem. Comparative analysis against conventional ETL-based architectures indicates potential structural advantages in storage efficiency (estimated 60-70% reduction), query freshness (near real-time versus batch-delayed), and governance coverage (automated versus manual lineage capture), while recognizing that end-to-end latency and throughput depend on source-system performance, network conditions, query complexity, and selective materialization for complex multi-source analytical workloads. The proposed LAIDW-MCP-DGF architecture offers a scalable, maintainable approach for organizations seeking to operationalize AI workflows with minimal data infrastructure overhead.

Monsinee Keeratikrainon · 0 citations
Open access 2024

Unified Metadata Management Framework for Hybrid Cloud Data Analytics

In the era of big data, hybrid cloud environments have become integral for data analytics, offering flexibility, scalability, and cost-efficiency. However, managing metadata across these complex, distributed systems remains a significant challenge. Metadata, which describes the data, its context, and its usage, plays a crucial role in improving data discovery, quality, and governance. In this paper, we propose a Unified Metadata Management Framework designed specifically for hybrid cloud data analytics. This framework integrates metadata from multiple cloud environments (public, private, and multi-cloud) into a cohesive, centralized system that facilitates efficient data management, governance, and analytics. We discuss the key components of the framework, including metadata repositories, metadata exchange layers, cataloging, and data lineage tracking. Additionally, the framework ensures security and privacy compliance while offering scalability and flexibility. We provide use cases from healthcare, finance, and e-commerce sectors to demonstrate its practical applications and evaluate its performance against traditional metadata management approaches. The proposed solution offers significant improvements in managing metadata across hybrid cloud infrastructures and lays the foundation for future innovations in cloud data analytics.

Ken Iverson · 0 citations
Open access 2023

Smart ERP: Scalable Data Engineering Frameworks Using Artificial Intelligence

The findings advocate for the integration of AI-powered pipelines within ERP systems as a transformative approach to enable scalable, intelligent, and high-fidelity data processing, essential for next- generation enterprise software resilience and performance.

Yuvaraj Kavala · 0 citations
Review Open access Aug 2026

Enterprise AI Transformation Through Modern Data Platforms

This review critically evaluates peer-reviewed journal literature published in the last decade related to AI capability, big data analytics capability, data governance, machine learning operations, digital transformation, and organizational value creation.

Saurabh Mishra · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.