Multi-Tenant Data Lake Architecture for Scalable AI and Big Data Workload Management
Abstract
The rapid expansion of artificial intelligence (AI), machine learning, computer vision, and multimodal analytics has increased the demand for data infrastructures capable of supporting heterogeneous workloads at large scale. Conventional data platforms frequently encounter difficulties when multiple users, applications, or organizational units simultaneously access shared datasets, compute resources, and analytical services. This paper develops a research-oriented conceptual architecture for a multi-tenant data lake designed to support scalable AI and big data workload management. The proposed architecture integrates tenant-aware data ingestion, metadata management, storage isolation, workload orchestration, resource governance, security, and adaptive AI processing into a unified framework. The methodology is derived through comparative synthesis of the supplied literature, including research on multimodal datasets, computer vision workloads, computational sciences, and responsible approaches to AI. The architecture emphasizes logical tenant isolation while preserving controlled opportunities for data and infrastructure sharing. The analysis indicates that workload-aware orchestration, metadata-driven resource allocation, and differentiated service policies can improve scalability and reduce resource contention in heterogeneous environments. The paper further argues that multi-tenancy must be treated not merely as a virtualization problem but as a data-governance, workload-management, and responsible-AI problem. The resulting framework provides a foundation for scalable AI data lakes while identifying limitations related to resource interference, governance complexity, data heterogeneity, and fairness.