Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-6· 0 citations· 9 references
Abstract
This paper presents a comprehensive analysis of edge AI architectures targeting embedded platforms and proposes a novel hierarchical design that addresses the critical challenges of computational efficiency, power consumption, and real-time processing in resource-constrained environments. The proposed architecture integrates adaptive quantization, dynamic load balancing, and multi-tier processing to optimize AI inference at the edge while maintaining high accuracy and low latency. Current edge AI implementations, such as ESP32-based systems, demonstrate the feasibility of bringing artificial intelligence to embedded devices, but lack the sophisticated resource management and scalability required for complex AI workloads. Our literature review reveals significant gaps in existing architectures, particularly in handling dynamic workloads and optimizing resource utilization across heterogeneous computing elements. We propose a three-tier hierarchical edge AI framework that couples adaptive mixedprecision quantization with a cross-tier load balancer and monitoring place, allowing the system to dynamically choose both precision and execution tier based on energy, latency, and accuracy constraints
The rapid growth of artificial intelligence (AI), machine learning, and large-scale digital services is placing unprecedented demand on cloud infrastructure, making scalable and sustainable compute increasingly important. While advances in processors and accelerators continue to improve computational capability, architectural coordination can unlock efficiency gains that hardware-generation improvements alone may not fully capture, particularly across increasingly heterogeneous compute environments (Armbrust et al., 2010). This paper will examine virtualization, workload scheduling, the provisioning of heterogeneous resources and infrastructure lifecycle management's effects on infrastructure efficiency at hyperscale. The paper extends the NIST cloud service definition (Mell & Grance, 2011) and most recent research on energy efficient resource management (Khan et al., 2022, Ilager et al., 2021) to propose a Sustainable Compute Lifecycle Framework comprised of five inter-connected phases: Platform Planning, Platform Deployment and Optimization, Platform Operations, Hardware Modernization, and Hardware Retirement and Resource Reclamation. The paper also explores new trends such as Compute Express Link (CXL) for memory pooling and disaggregation (Das Sharma et al., 2024;Chen et al., 2024) and the use of AI for infrastructure scheduling (Sanjalawe et al.,2025). The key message is that there is a new central design coordination of the cloud architecture that enables sustainable growth of the cloud, not merely increasing incremental capacity on hardware.
Priyadarshni Shanmugavadivelu· International journal of com...· 0 citations
A scalable edge-to-cloud AI inference pipeline in which inference tasks are dynamically distributed across heterogeneous edge and cloud resources is examined, providing a basis for resilient real-time AI systems while highlighting unresolved challenges involving heterogeneous hardware, dynamic workloads, privacy-utility trade-offs, and cross-layer optimization.
Khalid Al-Mansour· International Journal of Com...· 0 citations
Findings indicate that compression and knowledge distillation can reduce communication burdens, while heterogeneous aggregation and adaptive learning mechanisms improve the practicality of distributed AI environments.
Arif Setiawan, Maya Permata· International Journal of Com...· 0 citations
Enterprise adoption of artificial intelligence is restructuring the discipline of infrastructure planning in ways that conventional capacity models cannot accommodate. Artificial intelligence workloads span a heterogeneous spectrum of training, fine-tuning, inference, and batch scoring operations, each imposing qualitatively distinct demands on accelerator compute, storage throughput, and network fabric. The proliferation of graphics processing unit-accelerated clusters, high-bandwidth interconnects, and multi-cloud execution environments has rendered traditional provisioning frameworks inadequate for governing the scale, velocity, and compliance complexity inherent to production artificial intelligence platforms. This article presents a practitioner-oriented engineering framework for provisioning artificial intelligence-ready infrastructure that remains architecturally stable across accelerator generations, managed service evolutions, and organizational growth trajectories. Drawing on operational patterns from large-scale cloud transformation programs, the framework addresses workload segmentation, layered platform architecture, accelerator cluster governance, data provenance, network engineering, security, reliability, and cost governance as interdependent engineering concerns. The central argument is that organizations achieving sustained operational excellence in artificial intelligence infrastructure do so through deliberate platform architecture governed by automation-first operational practices, not through hardware procurement alone. The article concludes by projecting the long-term strategic implications of multi-cloud artificial intelligence transformation as a governed maturity progression, offering forward-looking guidance for infrastructure architects navigating an accelerating and mission-critical technology landscape
Hemanth Kumar Gandavarapu· International Journal of Eng...· 0 citations
A research-driven conceptual framework for resilient edge-to-cloud AI architectures supporting distributed real-time decision making and identifies limitations associated with heterogeneous devices, uncertain ground truth, model drift, communication failures, and the absence of uniform evaluation criteria are identified.
Chinedu Eze, F. Bello· International Journal of Adv...· 0 citations
The rapid expansion of artificial intelligence (AI), machine learning, computer vision, and multimodal analytics has increased the demand for data infrastructures capable of supporting heterogeneous workloads at large scale. Conventional data platforms frequently encounter difficulties when multiple users, applications, or organizational units simultaneously access shared datasets, compute resources, and analytical services. This paper develops a research-oriented conceptual architecture for a multi-tenant data lake designed to support scalable AI and big data workload management. The proposed architecture integrates tenant-aware data ingestion, metadata management, storage isolation, workload orchestration, resource governance, security, and adaptive AI processing into a unified framework. The methodology is derived through comparative synthesis of the supplied literature, including research on multimodal datasets, computer vision workloads, computational sciences, and responsible approaches to AI. The architecture emphasizes logical tenant isolation while preserving controlled opportunities for data and infrastructure sharing. The analysis indicates that workload-aware orchestration, metadata-driven resource allocation, and differentiated service policies can improve scalability and reduce resource contention in heterogeneous environments. The paper further argues that multi-tenancy must be treated not merely as a virtualization problem but as a data-governance, workload-management, and responsible-AI problem. The resulting framework provides a foundation for scalable AI data lakes while identifying limitations related to resource interference, governance complexity, data heterogeneity, and fairness.
Arjun Mehta, Priya Sharma· International Journal of Adv...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.