Skip to content

Towards Practical Latency SLOs on Cloud Data Warehouses

· 0 citations · 17 references

TL;DR

This work outlines AutoSLO, a latency-SLO-aware work-load management framework for multi-cluster cloud data ware-houses that includes a periodic Policy Tuner that proactively plans resources using workload fore-casts, an SLO-aware reactive Autoscaler that adjusts the active cluster set based on the observed workload, and an online Query Router that reacts to concurrent query load when routing.

View source

Similar papers

Open access May 2026

A Resource-centric Analysis and Optimization of NoSQL Workloads using Distressed Resource Volume Metric

This work proposes and develops an open-source policy simulation framework, LoadStar, which forms a reusable benchmark pipeline for validating policies for resource-centric NoSQL workloads, and defines a resource optimization problem for placing Cosmos DB replicas onto VM nodes, and develops the Luna model for forecasting future load distributions.

Gunika Verma, Aashutosh A, Pooja Srinivas et al. · 0 citations
Conference Jul 2026

SLO-Driven Horizontal Container Autoscaling

Modern web services are required to meet critical non-functional requirements, including availability, responsiveness, scalability, and reliability, which are formalized through Service Level Agreements (SLAs). SLAs define Service Level Objectives (SLOs), such as latency, throughput, and uptime, that ensure consistent service quality. Failing to meet these objectives can incur penalties and harm a provider's reputation. At the same time, over-provisioning resources leads to unnecessary costs and inefficient utilization. Autoscaling mechanisms address this by dynamically adjusting the number of service replicas according to demand. However, conventional approaches typically rely on low-level metrics, such as CPU or memory usage, which limit the ability to optimize both SLO compliance and infrastructure costs. This paper presents an enhanced SLO-driven autoscaling methodology for containerized workloads in Kubernetes clusters, integrating response time SLO targets into the autoscaling process. The proposed approach improves decision-making over traditional autoscaling by balancing service-level performance with operational efficiency. Experimental evaluation of a prototype demonstrates clear benefits compared to the default Kubernetes Horizontal Pod Autoscaler.

A. Marchese, O. Tomarchio · 0 citations
Open access Aug 2026

Multi-Tenant Data Lake Architecture for Scalable AI and Big Data Workload Management

The rapid expansion of artificial intelligence (AI), machine learning, computer vision, and multimodal analytics has increased the demand for data infrastructures capable of supporting heterogeneous workloads at large scale. Conventional data platforms frequently encounter difficulties when multiple users, applications, or organizational units simultaneously access shared datasets, compute resources, and analytical services. This paper develops a research-oriented conceptual architecture for a multi-tenant data lake designed to support scalable AI and big data workload management. The proposed architecture integrates tenant-aware data ingestion, metadata management, storage isolation, workload orchestration, resource governance, security, and adaptive AI processing into a unified framework. The methodology is derived through comparative synthesis of the supplied literature, including research on multimodal datasets, computer vision workloads, computational sciences, and responsible approaches to AI. The architecture emphasizes logical tenant isolation while preserving controlled opportunities for data and infrastructure sharing. The analysis indicates that workload-aware orchestration, metadata-driven resource allocation, and differentiated service policies can improve scalability and reduce resource contention in heterogeneous environments. The paper further argues that multi-tenancy must be treated not merely as a virtualization problem but as a data-governance, workload-management, and responsible-AI problem. The resulting framework provides a foundation for scalable AI data lakes while identifying limitations related to resource interference, governance complexity, data heterogeneity, and fairness.

Arjun Mehta, Priya Sharma · 0 citations
Open access 2026

A novel multi-authority access control scheme for fine grained access to users data in the cloud-based storage

Results demonstrate that HPA effectively responds to workload increases by provisioning additional pods, maintaining system stability and throughput during high-demand periods, with CPU usage and RPS exhibiting predictable scaling behavior aligned with the 15-second Metrics Server scraping interval.

Shamsuddeen Rabiu, Sani Muhammad Tanko, Eli Adama Jiya · 0 citations
Open access Aug 2026

Kernel-Level Dynamic Priority Scheduling for Containers

A dynamic priority scheduling framework at the kernel level that enhances the CPU allocation to latency-sensitive containers running in Kubernetes environments and reveals a significant improvement in terms of latency reduction, enhanced throughput, efficient utilization of CPU resources, and stable performance of scheduling under resource contention.

T. Rajkumar, Nishanth D., P. M et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.