Large Language Model-Assisted Metadata Engineering for Enterprise Data Platforms
Abstract
Enterprise metadata is essential for data discovery, governance, integration, and analytics. Traditional metadata engineering relies on manual, rule-based approaches that struggle with dynamic, heterogeneous enterprise data across cloud, IoT, ERP, CRM, and data lake environments. This study proposes a Large Language Model-Assisted Metadata Engineering Framework (LLM-MEF) that automates metadata extraction, semantic enrichment, schema recommendation, lineage discovery, and governance validation using transformer-based LLMs, Retrieval-Augmented Generation (RAG), vector databases, and knowledge graphs. The framework improves metadata quality, semantic consistency, discoverability, governance compliance, and operational efficiency while reducing manual effort. Explainable AI and continuous feedback learning further enhance transparency and adaptive improvement. The proposed LLM-MEF provides a scalable and intelligent solution for enterprise metadata management, supporting modern data governance, AI, business intelligence, regulatory compliance, and digital transformation.