Skip to content
Preprint

Optimizing API Gateway Placement in Multi-Cloud Kubernetes

Aug 2026 · 0 citations · 17 references
Computer Science

TL;DR

An optimization formulation that addresses API gateway placement as a capacitated facility location problem that jointly determines which candidate clusters to activate, how many gateway replicas to deploy, and how regional traffic should be distributed across the selected clusters is presented.

Abstract

The use of API gateways within geographically distributed multi-cloud Kubernetes clusters poses a tradeoff between infrastructure cost, computational resources, and network latencies. We present an optimization formulation that addresses API gateway placement as a capacitated facility location problem that jointly determines which candidate clusters to activate, how many gateway replicas to deploy, and how regional traffic should be distributed across the selected clusters. The formulation imposes an upper bound on estimated client-to-cluster network round-trip latency, excluding gateway processing, queuing, and backendservice latency, and incorporates a utilization headroom factor for gateway replica capacity. We present both a mixed-integer linear programming (MILP) formulation and a constructive greedy heuristic that ranks candidates according to incremental cost, comprising cluster-activation and marginal replica costs, per unit of assignable capacity while explicitly accounting for already committed load. Both formulations are applied to deterministic, seed-controlled, geography-based synthetic instances. For each problem size, 30 instances are generated with random seeds to analyze their performance. The greedy algorithm achieves an optimality gap of 3.2% to 4.7% to the MILP optimal solution, with a maximum observed gap of 25.0% for one particular instance, and a speedup of approximately 660x to 3,490x for 3 to 12 candidate clusters. In a canonical 10-candidate, 10-demand region instance, MILP-optimal deployment saves 24.2% in terms of monthly cost compared to the full-replication baseline. On the other hand, selecting the single cheapest candidate yields savings of 24.8% compared to the MILP optimum but does not satisfy the latency requirement for 3 out of 10 demand regions.

View source

Similar papers

Conference Aug 2026

MC-SPARK: a Policy-Constrained Genetic Algorithm for Multi-Cloud Workload Placement

Multi-cloud adoption expands pricing, performance, geographic, and resilience options, but makes workload placement combinatorial and policy dependent. This paper presents MC-SPARK, a policy-constrained genetic algorithm that maps workloads to candidate services while jointly considering operating cost, latency risk, a...

Krishnakumar Kunka Mohanram, Velu Natarajan, Arun Kumar Samayam · 0 citations
Preprint Sep 2026

Replication-Aware Placement of Functions and Data in the Edge-Cloud Continuum

This work introduces a Binary Linear Programming model to compute optimal placements and proposes a topology-aware greedy heuristic that efficiently approximates the optimal solution, making it suitable for periodic system reconfigurations.

Dario d'Abate, Matteo Cenzato, Matteo Briscini et al. · 0 citations
Open access Aug 2026

Janus: Joint Prefill/Decode Disaggregation with KV-Cache-Aware Multi-Cloud Routing for Edge-Adjacent LLM Serving

Janus is, to the authors' knowledge, the first scheduler to treat the KV-transport decision as a first-class scheduling variable jointly with prefill and decode routing across heterogeneous multi-cloud fleets, with provable guarantees.

K. Aruna, V. Kaliraj, I. Sudha et al. · 0 citations
#edge computing Preprint Sep 2026

Prediction-Robust Service Deployment with Capacity-Aware Edge Admission

This work proposes CAPSUM, a capacity-aware admission policy with an elastic specialization, CAPSUM-E, and implements an exact local offline dynamic program and compares against direct common-model baselines and documented source-derived adapters for EDP-A, OREO, and uEDC-L.

Hailiang Zhao, Zi-Qi Wang, Yi-Fei Zhang et al. · 0 citations
Preprint Aug 2026

A Kubernetes Scheduler Plugin for Cluster-Wide Placement Optimisation

The default scheduler of Kubernetes, the state-of-the-art container orchestrator, uses fast, local placement decisions. Unfortunately, this design leads to resource fragmentation, reduced cluster usage, and overprovisioning. External solvers can compute global placement plans, but enforcing these plans in upstream clus...

Henrik Christensen, Saverio Giallorenzo, J. Mauro · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.