Skip to content
Review Open access

Optimizing Cloud-Based Distributed Systems for Real-Time Machine Learning Model Deployment and Scaling

2021 · International Journal of Artificial Intelligence & Digital Transformation · 0 citations

TL;DR

Optimize techniques for cloud-based distributed systems to enhance the deployment and scaling of ML models in real-time applications and delves into the application of machine learning for cloud resource provisioning, emphasizing dynamic allocation based on real-time usage patterns.

Abstract

The rapid evolution of machine learning (ML) models and the surge in data volumes necessitate scalable and efficient deployment strategies. Cloud-based distributed systems offer on-demand scalability and resource flexibility, making them ideal for real-time ML model deployment and scaling. This paper explores optimization techniques for cloud-based distributed systems to enhance the deployment and scaling of ML models in real-time applications. We examine the integration of distributed systems and ML within cloud environments, focusing on scalable training and inference mechanisms. Key considerations such as task partitioning, communication overhead, fault tolerance, and resource optimization are discussed. Furthermore, we review auto-scaling techniques, highlighting advancements and challenges in dynamically adjusting resources to meet fluctuating demands. The paper also delves into the application of machine learning for cloud resource provisioning, emphasizing dynamic allocation based on real-time usage patterns. By synthesizing current research and practices, this study provides insights into effectively leveraging cloud-based distributed systems for real-time ML model deployment and scaling.

Read PDF

Similar papers

Open access 2025

Self-Adaptive Distributed Computing Models for High-Performance Analytics

This work proposes a scalable, intelligent, and resilient foundation for next-generation high-performance analytics and data-intensive applications that integrates adaptive resource management, intelligent workload scheduling, dynamic task migration, predictive analytics, and machine learning-based optimization to improve computational efficiency and responsiveness.

John Peterson, L. Martínez · 0 citations
Open access 2020

Enhancing Distributed Systems for Real-Time Machine Learning Model Deployment and Management

The integration of machine learning (ML) models into distributed systems has become pivotal for applications requiring real-time data processing and decision-making. This paper investigates methodologies to enhance distributed architectures for the efficient deployment and management of ML models in real-time environments. We explore the challenges associated with latency, scalability, and fault tolerance, and propose solutions leveraging edge computing, federated learning, and dynamic orchestration. Through empirical evaluations, we demonstrate the efficacy of the proposed approaches in optimizing real-time ML workflows.

S. Rahman · 0 citations
Open access 2024

Intelligent Workflow Scheduling for Distributed Data Processing Systems

An intelligent workflow scheduling framework that improves performance through adaptive decision-making, predictive analytics, and machine learning, and addresses key challenges like load balancing, scalability, energy efficiency, and fault tolerance is proposed.

D. Parnas · 0 citations
Open access 2024

AI-Driven Software Engineering: Optimizing Distributed Systems for Scalable Machine Learning Workflows

This paper explores the integration of artificial intelligence techniques into software engineering practices to optimize distributed systems for scalable machine learning (ML) workflows. As ML models grow in complexity and data volume, traditional system design approaches struggle to meet the demands of performance, scalability, and resource efficiency. We propose an AI-driven framework that leverages predictive analytics, automated resource management, and intelligent scheduling to enhance distributed computing environments. The study examines key challenges in distributed ML systems, including data partitioning, workload balancing, fault tolerance, and latency optimization. Through a combination of simulation and real-world case studies, we demonstrate how AI-based optimization strategies improve system throughput, reduce training time, and enhance resource utilization. The results highlight the potential of combining software engineering principles with AI-driven decision-making to build resilient and efficient ML infrastructures. This work contributes a structured approach for designing next-generation distributed systems capable of supporting large-scale machine learning applications.

Yuki Tanaka · 0 citations
Open access 2024

AI-Assisted Resource Scheduling in Multi-Cloud Computing Environments

This research proposes an AI-driven resource scheduling framework that integrates workload prediction, resource classification, intelligent scheduling, and continuous feedback mechanisms that aims to optimize multiple objectives, including cost reduction, execution efficiency, energy consumption, and SLA compliance.

Michael Anderson · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.