2023· International Journal of Data Engineering and Intelligent Computing· Vol 6, pp. 01-13· 0 citations
TL;DR
An ontology-guided approach to Data Mesh design is proposed, leveraging domain ontologies to ensure semantic interoperability, governance automation, and federated data product discovery and demonstrates how ontologies facilitate coherent data governance across distributed teams while preserving the autonomy of individual domains.
Abstract
The rise of Data Mesh as a paradigm for large-scale data management has challenged the traditional monolithic data lake approach by promoting decentralized ownership, domain-oriented architecture, and self-serve infrastructure. However, without a unified semantic layer, decentralization can lead to fragmentation, inconsistency, and governance bottlenecks. This paper proposes an ontology-guided approach to Data Mesh design, leveraging domain ontologies to ensure semantic interoperability, governance automation, and federated data product discovery. By aligning domain-specific knowledge structures with data product metadata, we demonstrate how ontologies facilitate coherent data governance across distributed teams while preserving the autonomy of individual domains. We explore architectural patterns, governance workflows, and implementation considerations, and present a case study to illustrate the application of this approach in a real-world enterprise setting.
In modern data-driven ecosystems, multi-domain data pipelines present significant challenges in terms of interoperability, scalability, governance, and semantic consistency. Traditional data integration methods often falter when faced with the heterogeneity and velocity of domain-specific data. This paper proposes a metadata-centric approach to managing the semantic layer in multi-domain data pipelines. By elevating metadata to a first-class citizen, organizations can dynamically model, interpret, and govern data semantics across domains without excessive data movement or manual schema alignment. We introduce a reference architecture and management framework that leverages active metadata, semantic annotations, knowledge graphs, and policy-driven governance to enable scalable, federated analytics. Case studies and experimental evaluation demonstrate the effectiveness of the approach in improving query performance, data lineage traceability, and cross-domain interoperability. The proposed methodology aligns with modern data mesh principles and promotes sustainable data infrastructure design.
K. Booth· International Journal of Dat...· 0 citations
In the evolving landscape of data architectures, the Data Mesh paradigm has emerged as a scalable and decentralized approach to managing complex data ecosystems. However, as data domains grow increasingly autonomous, ensuring consistency, discoverability, and transformation standardization across domains becomes challenging. This paper proposes a metadata-driven approach to data transformation within Data Mesh architectures. By leveraging active metadata such as data lineage, transformation rules, quality metrics, and semantic definitions data domains can automate and orchestrate transformations in a consistent, governed, and scalable manner. The paper explores architectural patterns, implementation considerations, and case studies demonstrating the benefits of this approach in enhancing interoperability, governance, and agility in enterprise data systems.
Lei Weing, Wei Chen· International Journal of Art...· 0 citations
Enterprise metadata is essential for data discovery, governance, integration, and analytics. Traditional metadata engineering relies on manual, rule-based approaches that struggle with dynamic, heterogeneous enterprise data across cloud, IoT, ERP, CRM, and data lake environments. This study proposes a Large Language Model-Assisted Metadata Engineering Framework (LLM-MEF) that automates metadata extraction, semantic enrichment, schema recommendation, lineage discovery, and governance validation using transformer-based LLMs, Retrieval-Augmented Generation (RAG), vector databases, and knowledge graphs. The framework improves metadata quality, semantic consistency, discoverability, governance compliance, and operational efficiency while reducing manual effort. Explainable AI and continuous feedback learning further enhance transparency and adaptive improvement. The proposed LLM-MEF provides a scalable and intelligent solution for enterprise metadata management, supporting modern data governance, AI, business intelligence, regulatory compliance, and digital transformation.
David Wheeler, Michael Gordon· International Journal of Dat...· 0 citations
In the era of digital transformation, organizations are increasingly leveraging multi-cloud strategies to harness the power of diverse cloud service providers. However, the heterogeneity of data schemas, storage formats, and semantic models across cloud platforms presents significant challenges for unified data analysis. Semantic data transformation offers a robust approach to harmonizing disparate datasets by mapping them to a common ontology or semantic model. This paper explores the methodologies, frameworks, and tools enabling semantic data transformation in multi-cloud data warehousing environments. It proposes an architecture for integrating semantic mediation, discusses best practices, and evaluates performance, scalability, and interoperability. Case studies and experimental results demonstrate the practical viability and impact of semantic transformation in achieving data consistency, governance, and advanced analytics across cloud platforms.
Ming Chen· International Journal of Dat...· 0 citations