Skip to content
Open access

A Scalable Multi-Tenant Framework for AI-Driven Big Data Lake Management and Processing

Aug 2026 · International Journal of Intelligent Data and Machine Learning · 0 citations

TL;DR

Analytical findings indicate that a policy-aware, AI-driven orchestration layer can improve workload prioritization, resource utilization, isolation, and operational transparency compared with static allocation models.

Abstract

The rapid growth of artificial intelligence (AI), machine learning, and heterogeneous data sources has increased the need for scalable data-lake architectures capable of supporting concurrent workloads, diverse data types, and dynamically changing computational requirements. Conventional data-management approaches often struggle to provide adequate isolation, resource elasticity, governance, and intelligent workload coordination in multi-tenant environments. This research proposes a conceptual scalable multi-tenant framework for AI-driven big data lake management and processing in which tenant-aware orchestration, workload classification, resource allocation, data governance, and AI-assisted decision mechanisms operate as integrated architectural components. The framework is theoretically grounded in scalable data-lake orchestration, formal reasoning, explainable decision-making, and resource-aware computational management. Particular emphasis is placed on separating tenant-level policies from shared infrastructure while maintaining efficient utilization of storage and processing resources. The architectural rationale is informed by research on rule-based learning, formal verification, satisfiability solving, explainability, and computational reasoning. The proposed framework further incorporates principles associated with multitenant data-lake orchestration for AI workloads (Goyal, 2025). Analytical findings indicate that a policy-aware, AI-driven orchestration layer can improve workload prioritization, resource utilization, isolation, and operational transparency compared with static allocation models. The study also identifies limitations associated with governance complexity, model dependence, computational overhead, and fairness across tenants. The resulting framework provides a research foundation for scalable, intelligent, and explainable big data lake management.

Read PDF

Similar papers

Open access Aug 2026

Multi-Tenant Data Lake Architecture for Scalable AI and Big Data Workload Management

The rapid expansion of artificial intelligence (AI), machine learning, computer vision, and multimodal analytics has increased the demand for data infrastructures capable of supporting heterogeneous workloads at large scale. Conventional data platforms frequently encounter difficulties when multiple users, applications, or organizational units simultaneously access shared datasets, compute resources, and analytical services. This paper develops a research-oriented conceptual architecture for a multi-tenant data lake designed to support scalable AI and big data workload management. The proposed architecture integrates tenant-aware data ingestion, metadata management, storage isolation, workload orchestration, resource governance, security, and adaptive AI processing into a unified framework. The methodology is derived through comparative synthesis of the supplied literature, including research on multimodal datasets, computer vision workloads, computational sciences, and responsible approaches to AI. The architecture emphasizes logical tenant isolation while preserving controlled opportunities for data and infrastructure sharing. The analysis indicates that workload-aware orchestration, metadata-driven resource allocation, and differentiated service policies can improve scalability and reduce resource contention in heterogeneous environments. The paper further argues that multi-tenancy must be treated not merely as a virtualization problem but as a data-governance, workload-management, and responsible-AI problem. The resulting framework provides a foundation for scalable AI data lakes while identifying limitations related to resource interference, governance complexity, data heterogeneity, and fairness.

Arjun Mehta, Priya Sharma · 0 citations
Aug 2026

AI-Enabled Data Lake Architecture for Efficient Multi-Tenant Big Data Orchestration

The analysis indicates that combining adaptive exploration with value-based decision mechanisms can provide a stronger orchestration model than static policies, although computational overhead, training instability, tenant fairness, and limited empirical validation remain important constraints.

Faisal Alharbi, Sara Al-Qahtani · 0 citations
Open access Aug 2026

Adaptive Multi-Agent AI Framework for Real-Time Data Streaming with Enhanced Scalability and Resilience

Real-time data streaming systems increasingly operate under highly variable workloads, heterogeneous data sources, latency constraints, and frequent service disruptions. Conventional stream-processing architectures generally depend on predefined routing, static resource allocation, and centralized coordination, which can limit their ability to adapt when event rates, computational requirements, or infrastructure conditions change rapidly. This paper proposes an Adaptive Multi-Agent AI Framework for Real-Time Data Streaming with Enhanced Scalability and Resilience, in which autonomous AI agents collaboratively perform stream monitoring, workload classification, task allocation, resource adaptation, anomaly detection, and recovery. The theoretical foundation combines multi-agent coordination with contextual representation, long-document processing, memory management, and adaptive decision-making. Prior work on aspect-controllable summarization demonstrates the value of controlling computational objectives according to task requirements, while studies of coreference, lexical chains, and entity-based coherence emphasize the importance of preserving relationships across distributed information units (Amplayo, Angelidis, & Lapata, 2021; Baldwin & Morton, 1998; Barzilay & Elhadad, 1997; Barzilay & Lapata, 2005). Long-context language modeling further motivates mechanisms capable of retaining relevant information over extended streaming windows (Beltagy, Peters, & Cohan, 2020). The proposed framework extends these principles to adaptive streaming environments and aligns with recent multi-agent event-streaming research emphasizing resiliency and scalability (Reddy et al., 2026). Analytical findings indicate that decentralized agent specialization, shared contextual state, adaptive workload redistribution, and failure-aware coordination can provide a stronger basis for resilient streaming than static pipelines. The paper also identifies trade-offs involving coordination overhead, state consistency, model complexity, and resource consumption.

Nethmi Perera, K. Fernando · 0 citations
Open access 2025

Autonomous Data Fabric Architectures for Enterprise-Wide Intelligent Computing

Modern enterprises generate massive volumes of data from cloud platforms, IoT devices, enterprise applications, social media, and AI systems, creating challenges in data integration, governance, scalability, security, and real-time analytics. Traditional data management approaches often struggle to handle these complex and distributed environments. This paper proposes an Autonomous Data Fabric (ADF) architecture that combines AI/ML, metadata-driven automation, knowledge graphs, intelligent orchestration, and policy-based governance to enable seamless, self-managing enterprise data ecosystems. The framework supports automated data discovery, semantic integration, adaptive workflows, continuous monitoring, and intelligent resource optimization while ensuring data quality, security, and compliance. Experimental results demonstrate that the proposed ADF significantly improves data integration efficiency, governance, analytics performance, operational cost, and decision-making compared to conventional systems. Its scalable and self-adaptive design supports hybrid cloud, multi-cloud, edge, and on-premises environments, making it a robust solution for enterprise digital transformation and next-generation intelligent data management.

Narendra Karmarkar · 0 citations
Open access 2024

AI-Assisted Data Pipeline Orchestration for Scalable Analytics

Modern enterprises face increasing demands for scalable and efficient data processing due to rapid data growth. Traditional data pipeline orchestration methods, which rely on static configurations and manual intervention, often lead to inefficiencies in resource use, latency, and fault tolerance. This paper proposes an AI-assisted orchestration framework that integrates machine learning techniques to enable dynamic scheduling, workload prediction, anomaly detection, and resource optimization. By leveraging reinforcement learning, supervised learning, and heuristic methods, the system adapts pipeline configurations in real time based on changing workloads and system conditions. The proposed architecture includes data ingestion modules, AI-driven orchestration engines, adaptive schedulers, and monitoring systems. A key contribution is an intelligent scheduling mechanism that improves execution efficiency and resource utilization. Experimental results show significant improvements over traditional systems, with up to 35% increase in processing efficiency and 25% reduction in latency. The study concludes that AI-driven orchestration is a promising approach for building scalable and autonomous data processing systems, with future work focusing on deeper integration of advanced learning models and real-time adaptability.

J. Weizenbaum, S. Papert · 1 citation
Open access 2023

Smart ERP: Scalable Data Engineering Frameworks Using Artificial Intelligence

The findings advocate for the integration of AI-powered pipelines within ERP systems as a transformative approach to enable scalable, intelligent, and high-fidelity data processing, essential for next- generation enterprise software resilience and performance.

Yuvaraj Kavala · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.