A three-layer architecture that addresses the tension between elastic, multi-tenant cloud operations and performance-critical HPC batch workflows, and places cloud-style Infrastructure-as-a-Service (IaaS) management, based on virtualization with GPU passthrough, at the center.
Abstract
The rapid expansion of artificial intelligence (AI) workloads is prompting governments and institutions worldwide to invest in sovereign AI Factories: large-scale infrastructures designed to support both cloud-native AI services and AI-oriented HPC workloads. This paper proposes a three-layer architecture that addresses the tension between elastic, multi-tenant cloud operations and performance-critical HPC batch workflows. We position the Cloud Layer as the central orchestration and resource management plane, responsible for provisioning GPU-accelerated virtual machines with near-bare-metal throughput and for coordinating compute, GPU, network, and storage resources under multi-tenant isolation. Unlike Kubernetes-centric designs, our architecture places cloud-style Infrastructure-as-a-Service (IaaS) management, based on virtualization with GPU passthrough, at the center. This provides strong multi-tenancy, clearer security boundaries, and hardware-level isolation between tenants, while remaining agnostic to the workload orchestrator. We ground the proposal in a concrete reference implementation: the OpenNebula AI Factory Reference Architecture, which is used in large-scale European initiatives such as IPCEI-CIS. Experimental evaluation with the cuBLAS benchmark confirms that GPU passthrough virtualization introduces negligible performance overhead compared with bare metal across three representative GPU generations (NVIDIA GB200, H100, and L40S), including configurations using MIG partitioning, for compute-bound workloads. We analyze key challenges in resource allocation, placement, and partitioning, and discuss design trade-offs and open research directions for efficient, sovereign AI infrastructures.
Enterprise adoption of artificial intelligence is restructuring the discipline of infrastructure planning in ways that conventional capacity models cannot accommodate. Artificial intelligence workloads span a heterogeneous spectrum of training, fine-tuning, inference, and batch scoring operations, each imposing qualitatively distinct demands on accelerator compute, storage throughput, and network fabric. The proliferation of graphics processing unit-accelerated clusters, high-bandwidth interconnects, and multi-cloud execution environments has rendered traditional provisioning frameworks inadequate for governing the scale, velocity, and compliance complexity inherent to production artificial intelligence platforms. This article presents a practitioner-oriented engineering framework for provisioning artificial intelligence-ready infrastructure that remains architecturally stable across accelerator generations, managed service evolutions, and organizational growth trajectories. Drawing on operational patterns from large-scale cloud transformation programs, the framework addresses workload segmentation, layered platform architecture, accelerator cluster governance, data provenance, network engineering, security, reliability, and cost governance as interdependent engineering concerns. The central argument is that organizations achieving sustained operational excellence in artificial intelligence infrastructure do so through deliberate platform architecture governed by automation-first operational practices, not through hardware procurement alone. The article concludes by projecting the long-term strategic implications of multi-cloud artificial intelligence transformation as a governed maturity progression, offering forward-looking guidance for infrastructure architects navigating an accelerating and mission-critical technology landscape
Hemanth Kumar Gandavarapu· International Journal of Eng...· 0 citations
The rapid growth of artificial intelligence (AI), machine learning, and large-scale digital services is placing unprecedented demand on cloud infrastructure, making scalable and sustainable compute increasingly important. While advances in processors and accelerators continue to improve computational capability, architectural coordination can unlock efficiency gains that hardware-generation improvements alone may not fully capture, particularly across increasingly heterogeneous compute environments (Armbrust et al., 2010). This paper will examine virtualization, workload scheduling, the provisioning of heterogeneous resources and infrastructure lifecycle management's effects on infrastructure efficiency at hyperscale. The paper extends the NIST cloud service definition (Mell & Grance, 2011) and most recent research on energy efficient resource management (Khan et al., 2022, Ilager et al., 2021) to propose a Sustainable Compute Lifecycle Framework comprised of five inter-connected phases: Platform Planning, Platform Deployment and Optimization, Platform Operations, Hardware Modernization, and Hardware Retirement and Resource Reclamation. The paper also explores new trends such as Compute Express Link (CXL) for memory pooling and disaggregation (Das Sharma et al., 2024;Chen et al., 2024) and the use of AI for infrastructure scheduling (Sanjalawe et al.,2025). The key message is that there is a new central design coordination of the cloud architecture that enables sustainable growth of the cloud, not merely increasing incremental capacity on hardware.
Priyadarshni Shanmugavadivelu· International journal of com...· 0 citations
This paper explores hybrid cloud-edge infrastructures as a scalable solution for deploying AI in IIoT environments and presents an architectural framework that balances compute-intensive model training in the cloud with low-latency inference at the edge.
Jennifer Clark· International Journal of Mac...· 0 citations
The rapid expansion of artificial intelligence (AI), machine learning, computer vision, and multimodal analytics has increased the demand for data infrastructures capable of supporting heterogeneous workloads at large scale. Conventional data platforms frequently encounter difficulties when multiple users, applications, or organizational units simultaneously access shared datasets, compute resources, and analytical services. This paper develops a research-oriented conceptual architecture for a multi-tenant data lake designed to support scalable AI and big data workload management. The proposed architecture integrates tenant-aware data ingestion, metadata management, storage isolation, workload orchestration, resource governance, security, and adaptive AI processing into a unified framework. The methodology is derived through comparative synthesis of the supplied literature, including research on multimodal datasets, computer vision workloads, computational sciences, and responsible approaches to AI. The architecture emphasizes logical tenant isolation while preserving controlled opportunities for data and infrastructure sharing. The analysis indicates that workload-aware orchestration, metadata-driven resource allocation, and differentiated service policies can improve scalability and reduce resource contention in heterogeneous environments. The paper further argues that multi-tenancy must be treated not merely as a virtualization problem but as a data-governance, workload-management, and responsible-AI problem. The resulting framework provides a foundation for scalable AI data lakes while identifying limitations related to resource interference, governance complexity, data heterogeneity, and fairness.
Arjun Mehta, Priya Sharma· International Journal of Adv...· 0 citations
An integrated management framework powered by Tapis is shown how a unified UI-driven workflow streamlines the transition from initial evaluation to deployment, ensuring operational consistency and reproducibility without manual script porting.
Manikya Swathi Vallabhajosyula, Gautam Gururaj Molakalmuru, Samuel Khuvis et al.· Practice and Experience in A...· 0 citations
An architecture that decouples complex policy enforcement from high-speed packet forwarding to support VPC semantics on back-end NICs and enable front-end/back-end integration is proposed, suggesting that commodity hardware can support both high-throughput AI training and flexible VPC features.
Yinhe Wang, Xing Li, Enge Song et al.· Asia-Pacific Workshop on Net...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.