Skip to content

PREreview of "You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference"

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Cloud Computing and Resource Management

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23147092. ## Summary This paper identifies a genuinely under-appreciated routing axis: after a model router picks Llama-3.3-70B, the client must still choose *which provider serves it* — and the price list is a bad guide. Measuring live endpoints across 6 open models, multiple providers, and three waves over 43 days, the authors find price predicts latency (median Spearman −0.61) but not accuracy (+0.05) or availability (0.00); feasibility is task-selective (one Llama-3.3-70B deployment is near-normal on knowledge tasks but catastrophically degraded on multi-step reasoning); and the map drifts (routes changed in 4/17 comparable cells over days 0–13). Their measured-map policy — cheapest provider within 5 points of best measured accuracy with >90% availability — yields median 50% savings vs. the premium provider at matched quality. FACET, their online router, certifies per-(provider×task) feasibility facets before serving cheap endpoints, fails safe to an anchor, and monitors certified facets with a slip detector. A 36-hour live run (Llama-3.3-70B, 12 providers, 538 queries) cut average served price from $0.924/M to $0.210/M tokens — 63.7% serving-cost reduction (57.1% incl. probes) at 88.3% accuracy. ## Strengths 1. **The decision axis is real and well-positioned.** Same weights, different quantization/kernels/batching, wildly different service — anyone running multi-provider open-weight inference has felt this. Appendix B's Table 7 draws the line cleanly against FrugalGPT/RouteLLM/CARROT/MixLLM: those pick the model; this picks the server underneath, and the two layers compose (Section 5). 2. **Unusually honest measurement reporting.** Three-wave drift study, thorough Appendix A limitations (attribution-agnostic claims; pinned probing can induce rate limits; single aggregator and client location; price–latency is associational, not causal). The task-selective failure finding alone is worth the paper. 3. **FACET is framed as insurance, not optimization.** Separating certification (admission) from monitoring (drift) with component ablations (Tables 3–4), plus an explicit break-even analysis of the safety premium (~17× an ordinary query vs. SW-UCB, Appendix E), is the right framing: you pay for the anchor to avoid serving a catastrophic mine. ## Major concerns 1. **The live deployment never tests the core mechanism.** The single 36-hour run is caveated: "no provider quality collapse occurred during this window, so the experiment validates cold-start certification and live migration rather than slip detection." Slip detection is FACET's raison d'être, and it is exercised only in exact-replay — with Appendix A admitting the ablation and bias stress tests all use a single Llama-3.3-70B/GSM8K replay cell. The headline experiment should show the detector catching a real slip, on multiple cells, live. 2. **Threshold fragility is buried in the appendices.** The 50% saving rests on δ=0.05, nmin=12, 90% availability — but the out-of-sample below-floor rate swings from 0% to 22.2% across threshold choices (Appendix H), and the confidence-based analysis cuts median savings from 55% to 46% (Appendix U). The main text should carry uncertainty-quantified savings, not point-estimate victory margins. 3. **The trust boundary reintroduces the cost the paper claims to remove.** Systematic evaluator bias "can corrupt certification unless ground-truth probes or audits provide an independent quality signal" — but ground-truth labels for arbitrary production queries are exactly what doesn't exist. And the anchor must be maintained: in practice it's the premium provider you were trying to escape. Who pays for the anchor and the gold audits is left unquantified. ## Practitioner perspective The task-selective failure finding is the one I'd pin to every team's wall: provider safety is a (provider × task × time) measurement, not a label. But the production math that matters — anchor cost + probe cost + gold-audit cost vs. measured-map saving — is only half quantified. Before adopting FACET, I'd want the slip-detection demo on live traffic and the full cost of the trust infrastructure in one table. ## Overall assessment An important paper with novel framing, careful measurement, and rare honesty about limitations. Recommend with revisions: demonstrate slip detection live (or across multiple replay cells), move uncertainty-quantified savings into the main text, and quantify who pays for the anchor and ground-truth probes. Competing interests The author declares that they have no competing interests. Use of Artificial Intelligence (AI) The author declares that they used generative AI to come up with new ideas for their review.

View source

Similar papers

#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations...

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17
#artificial intelligence Open access Jul 2024

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

This work investigates the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection and finds that the MTEs are biased, significantly favoring White-associated names in 85% of cases and female-associated names in only 11.1% of cases.

Kyra Wilson, Aylin Caliskan · 131 citations · ⚡8
#artificial intelligence Review Oct 2025

Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

This paper presents a comprehensive overview of the Ultralytics YOLO family, emphasizing architectural evolution, benchmarking, deployment, and emerging directions from YOLOv5 through YOLO27, and examines detection, segmentation, depth, classification, pose, oriented detection, tracking, export, quantization, and deplo...

Ranjan Sapkota, Manoj Karkee · 112 citations · ⚡10

BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models

A novel threat is unveiled in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base, enabling the attacker to steer the response without altering the user input or modifying the RAG weights.

Jiaqi Xue, Meng Zheng, Yebowen Hu et al. · 109 citations · ⚡8

The Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models

This work revisits schema linking when using the latest generation of large language models (LLMs) and finds empirically that newer models are adept at utilizing relevant schema elements during generation even in the presence of large numbers of irrelevant ones.

Karime Maamari, Fadhil Abubaker, Daniel Jaroslawicz et al. · 109 citations · ⚡19

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.

Jinhe Bi, Yifan Wang, Danqi Yan et al. · 73 citations · ⚡4

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.