Skip to content
Book Open access

Taming Inference Workloads at Global Scale: Foundation Model Serving in Amazon Bedrock

Sep 2026 · Proceedings of the ACM SIGOPS 32nd Symposium on Operating Systems Principles · 0 citations · 27 references

Abstract

Foundation model (FM) inference platforms must manage scarce accelerator capacity distributed unevenly across regions. They must also handle requests whose token consumption varies widely and may be revealed progressively during generation, while sharing capacity across workloads with different latency and throughput objectives. This paper describes the architecture of the Workload Management System (WLM) deployed in Amazon Bedrock, which serves millions of inference requests per minute across more than 30 regions. The architecture follows two principles. First, decouple request ingress from execution to form a logical execution pool within permitted geographic and model-placement scopes, with prompt-cache stability as a first-class routing constraint. Second, favor bounded inconsistency with defense in depth over synchronous global coordination. WLM comprises five cooperating subsystems: admission control, traffic classification, cross-region routing, request scheduling, and orchestration. Admission and routing decisions use local or cached state; quota consumption is reconciled periodically across regions, while routing weights are refreshed off the request path from regional capacity and health signals. We evaluate WLM using production observations, benchmark measurements, and controlled load tests comparing bounded-staleness with synchronous admission. The evaluation characterizes accelerator utilization, cache stability, request latency, priority behavior, throughput, request-path overhead, quota-enforcement accuracy, and behavior under selected failures.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.