Skip to content
Open access

Scalable AI Inference Pipelines Across Edge and Cloud Computing Environments

Khalid Al-Mansour
Aug 2026 · International Journal of Computer Science & Information System · 0 citations

TL;DR

A scalable edge-to-cloud AI inference pipeline in which inference tasks are dynamically distributed across heterogeneous edge and cloud resources is examined, providing a basis for resilient real-time AI systems while highlighting unresolved challenges involving heterogeneous hardware, dynamic workloads, privacy-utility trade-offs, and cross-layer optimization.

Abstract

The increasing deployment of artificial intelligence (AI) applications in healthcare, industrial Internet of Things (IIoT), intelligent transportation, and next-generation wireless systems has created a demand for inference architectures that simultaneously provide low latency, scalability, privacy, reliability, and efficient resource utilization. Conventional cloud-centric inference architectures provide substantial computational capacity but can introduce network latency, bandwidth consumption, privacy exposure, and dependence on centralized infrastructure. Edge computing addresses several of these limitations by relocating computation closer to data sources, while cloud environments remain important for computationally intensive and globally coordinated workloads. This research examines a scalable edge-to-cloud AI inference pipeline in which inference tasks are dynamically distributed across heterogeneous edge and cloud resources. The methodology synthesizes the provided literature on federated learning, edge resource allocation, dynamic scheduling, privacy preservation, machine learning for 6G, IIoT, and secure healthcare systems. A layered architectural model is developed around workload characterization, adaptive task placement, communication-aware scheduling, privacy protection, and resilient orchestration. The analysis indicates that scalability is not achieved merely by adding computational resources; rather, it depends on coordinated optimization of computation, communication, privacy, and scheduling. The proposed conceptual framework positions edge inference as the first computational layer, cloud inference as an elastic computational layer, and intelligent orchestration as the mechanism connecting the two. The resulting architecture provides a basis for resilient real-time AI systems while highlighting unresolved challenges involving heterogeneous hardware, dynamic workloads, privacy-utility trade-offs, and cross-layer optimization.

Read PDF

Similar papers

Review Open access Aug 2026

Edge-Cloud AI Computing: A Robust Framework for Real-Time Inference and Decision Automation

Findings indicate that compression and knowledge distillation can reduce communication burdens, while heterogeneous aggregation and adaptive learning mechanisms improve the practicality of distributed AI environments.

Arif Setiawan, Maya Permata · 0 citations
Open access Aug 2026

Resilient Edge-to-Cloud AI Architectures for Distributed Real-Time Decision Making

A research-driven conceptual framework for resilient edge-to-cloud AI architectures supporting distributed real-time decision making and identifies limitations associated with heterogeneous devices, uncertain ground truth, model drift, communication failures, and the absence of uniform evaluation criteria are identified.

Chinedu Eze, F. Bello · 0 citations
Conference Open access Jul 2026

Profiling Neural Network Partitioning Strategies for Inference across the Computing Continuum

A formal model that unifies functional, pipelined, and data-parallel partitioning strategies within a single abstraction over heterogeneous CC topologies, enabling structured cross-strategy comparison and enabling empirical, cross-strategy comparison of distributed inference deployments is introduced.

Nikolaos Papadakis, Alexandros Angourakis, K. Magoutis et al. · 0 citations
Open access 2020

Hybrid Cloud-Edge Infrastructures for Scalable IIoT AI Deployments

This paper explores hybrid cloud-edge infrastructures as a scalable solution for deploying AI in IIoT environments and presents an architectural framework that balances compute-intensive model training in the cloud with low-latency inference at the edge.

Jennifer Clark · 0 citations
Conference Jul 2026

Novel Hierarchical Edge AI Architecture for Resource-Constrained Embedded Platforms: A Comprehensive Framework for Distributed Intelligence

This paper presents a comprehensive analysis of edge AI architectures targeting embedded platforms and proposes a novel hierarchical design that addresses the critical challenges of computational efficiency, power consumption, and real-time processing in resource-constrained environments. The proposed architecture integrates adaptive quantization, dynamic load balancing, and multi-tier processing to optimize AI inference at the edge while maintaining high accuracy and low latency. Current edge AI implementations, such as ESP32-based systems, demonstrate the feasibility of bringing artificial intelligence to embedded devices, but lack the sophisticated resource management and scalability required for complex AI workloads. Our literature review reveals significant gaps in existing architectures, particularly in handling dynamic workloads and optimizing resource utilization across heterogeneous computing elements. We propose a three-tier hierarchical edge AI framework that couples adaptive mixedprecision quantization with a cross-tier load balancer and monitoring place, allowing the system to dynamically choose both precision and execution tier based on energy, latency, and accuracy constraints

Ravi Suppiah, Ravichandran Danthakanni, M. Nair et al. · 0 citations
Review Open access Aug 2026

A Intelligent Edge-Cloud Integration for Resilient and Real-Time AI Decision Systems

The analysis indicates that effective edge-cloud AI systems require adaptive workload placement, privacy-preserving distributed learning, security-aware inference, explainability, fault tolerance, and continuous resource optimization rather than simple physical distribution of computation.

Amir Hosseini, L. Karimi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.