Scalable AI Inference Pipelines Across Edge and Cloud Computing Environments
Khalid Al-Mansour
Aug 2026· International Journal of Computer Science & Information System· 0 citations
TL;DR
A scalable edge-to-cloud AI inference pipeline in which inference tasks are dynamically distributed across heterogeneous edge and cloud resources is examined, providing a basis for resilient real-time AI systems while highlighting unresolved challenges involving heterogeneous hardware, dynamic workloads, privacy-utility trade-offs, and cross-layer optimization.
Abstract
The increasing deployment of artificial intelligence (AI) applications in healthcare, industrial Internet of Things (IIoT), intelligent transportation, and next-generation wireless systems has created a demand for inference architectures that simultaneously provide low latency, scalability, privacy, reliability, and efficient resource utilization. Conventional cloud-centric inference architectures provide substantial computational capacity but can introduce network latency, bandwidth consumption, privacy exposure, and dependence on centralized infrastructure. Edge computing addresses several of these limitations by relocating computation closer to data sources, while cloud environments remain important for computationally intensive and globally coordinated workloads. This research examines a scalable edge-to-cloud AI inference pipeline in which inference tasks are dynamically distributed across heterogeneous edge and cloud resources. The methodology synthesizes the provided literature on federated learning, edge resource allocation, dynamic scheduling, privacy preservation, machine learning for 6G, IIoT, and secure healthcare systems. A layered architectural model is developed around workload characterization, adaptive task placement, communication-aware scheduling, privacy protection, and resilient orchestration. The analysis indicates that scalability is not achieved merely by adding computational resources; rather, it depends on coordinated optimization of computation, communication, privacy, and scheduling. The proposed conceptual framework positions edge inference as the first computational layer, cloud inference as an elastic computational layer, and intelligent orchestration as the mechanism connecting the two. The resulting architecture provides a basis for resilient real-time AI systems while highlighting unresolved challenges involving heterogeneous hardware, dynamic workloads, privacy-utility trade-offs, and cross-layer optimization.
Findings indicate that compression and knowledge distillation can reduce communication burdens, while heterogeneous aggregation and adaptive learning mechanisms improve the practicality of distributed AI environments.
Arif Setiawan, Maya Permata· International Journal of Com...· 0 citations
A research-driven conceptual framework for resilient edge-to-cloud AI architectures supporting distributed real-time decision making and identifies limitations associated with heterogeneous devices, uncertain ground truth, model drift, communication failures, and the absence of uniform evaluation criteria are identified.
Chinedu Eze, F. Bello· International Journal of Adv...· 0 citations
A formal model that unifies functional, pipelined, and data-parallel partitioning strategies within a single abstraction over heterogeneous CC topologies, enabling structured cross-strategy comparison and enabling empirical, cross-strategy comparison of distributed inference deployments is introduced.
Nikolaos Papadakis, Alexandros Angourakis, K. Magoutis et al.· International Conference on...· 0 citations
This paper explores hybrid cloud-edge infrastructures as a scalable solution for deploying AI in IIoT environments and presents an architectural framework that balances compute-intensive model training in the cloud with low-latency inference at the edge.
Jennifer Clark· International Journal of Mac...· 0 citations
This paper presents a comprehensive analysis of edge AI architectures targeting embedded platforms and proposes a novel hierarchical design that addresses the critical challenges of computational efficiency, power consumption, and real-time processing in resource-constrained environments. The proposed architecture integrates adaptive quantization, dynamic load balancing, and multi-tier processing to optimize AI inference at the edge while maintaining high accuracy and low latency. Current edge AI implementations, such as ESP32-based systems, demonstrate the feasibility of bringing artificial intelligence to embedded devices, but lack the sophisticated resource management and scalability required for complex AI workloads. Our literature review reveals significant gaps in existing architectures, particularly in handling dynamic workloads and optimizing resource utilization across heterogeneous computing elements. We propose a three-tier hierarchical edge AI framework that couples adaptive mixedprecision quantization with a cross-tier load balancer and monitoring place, allowing the system to dynamically choose both precision and execution tier based on energy, latency, and accuracy constraints
Ravi Suppiah, Ravichandran Danthakanni, M. Nair et al.· 2026 6th International Confe...· 0 citations
The analysis indicates that effective edge-cloud AI systems require adaptive workload placement, privacy-preserving distributed learning, security-aware inference, explainability, fault tolerance, and continuous resource optimization rather than simple physical distribution of computation.
Amir Hosseini, L. Karimi· International Journal of Adv...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.