Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads
A lightweight latency estimator is designed that simulates autoregressive token batch processing in the serving framework and estimates the time-to-first-token (TTFT) of queries and is incorporated into a latency-aware router that jointly optimizes latency, accuracy, and cost when assigning queries to model instances.