Skip to content
#small language model Open access

PREreview of "Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving"

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Energy Efficiency in Computing Green IT and Sustainability

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23177620. ## Summary Routing papers usually optimize proxy costs — parameter counts, API prices, model counts. This one measures the thing itself: per-query GPU energy. The authors run an offline tournament over a fixed pool of small-to-mid open models, recording correctness, latency, board power (polled via pynvml at 200ms, idle baseline subtracted), and effective GPU energy for each query. A small LLM router is then trained on these measurements — supervised fine-tuning followed by GRPO with an energy coefficient λe in the reward — to select one answer model per query. On 913 held-out questions across seven benchmarks, the KL-GRPO checkpoint matches UniRoute (60.9% vs 61.8% accuracy) at ~23% lower mean answer energy (1.69 vs 2.19 kJ); a GRPO checkpoint sits near Smoothie (58.4% vs 59.9%) at ~24% lower energy. The paper also reports a sharp accuracy–energy phase transition across routers, and a striking training effect: RL cuts the controller's own decode from ~2.4s to ~0.8s, making the router itself cheaper. Training cost is quantified as a one-time 0.32–0.40 kWh. ## Strengths 1. **Measured joules, not proxies — the right measurement discipline.** Energy is the cost dimension routing research has mostly ignored, and per-query pynvml profiling with idle-baseline subtraction is how you do it honestly. The finding that MATH-500 is 13% of questions but 52–57% of energy is exactly the kind of measurement that changes how a practitioner thinks about their workload mix. 2. **Admirably honest about failure modes.** The paper reports policy collapse outright — GRPO at λe=0.3 routes literally everything to Llama 3.1 8B with zero selection entropy — plus SFT over-selecting the 0.5B model, and the mid-size-arm aggregate illusion (Always-Qwen2.5-7B looks strong at 64.0% overall but collapses to 26.2% on BBH). These are the results that make the rest of the paper believable. 3. **RL made the router itself cheaper — a genuinely interesting effect.** The 2.4s→0.8s controller decode improvement means RL didn't just pick cheaper answers, it learned to decide faster. Router overhead is the perennial objection to learned routing, and this is the first paper I've seen that measures the router's own cost shrinking under training. ## Major concerns 1. **Offline lookup on one A100 is not serving.** All evaluation is offline replay of pre-logged tournament outcomes — no live serving, no batching, no continuous batching, no multi-tenancy, one A100-SXM4-80GB via Ollama. Production energy per query depends on batching, PUE, idle-power amortization, and hardware mix; joules measured in a lab rig don't translate to anyone's bill without those. The paper's own future work admits online and multi-turn agent settings are unevaluated — that's where the energy actually lives. 2. **The controller's energy is excluded from the headline comparison.** The authors state they "do not invent a controller cost" for baselines — fair — but the main tables then report answer-model energy only, while their own Figure 5 shows controller energy in the hundreds of joules. On cheap queries the router can cost nearly as much as the answer it selects; the headline 23% saving is computed on a ledger that omits the router. 3. **The pool excludes the models where routing actually saves money.** Everything is 0.5B–32B open models; no frontier or API-billed models, and the accuracy band tops out around 60–62% while always-largest hits 73.5%. The router never reaches just-use-the-big-model accuracy, so the "savings" are measured inside a pool nobody routes between in production — the real routing decision (small vs. frontier API) is outside the experiment entirely. ## Practitioner perspective I'd buy the measurement methodology before the routing policy: run the offline tournament on your own fleet, in joules, and you'll learn more about your workload than any router paper will tell you. Energy-aware routing matters most for on-prem and GPU-constrained serving, where joules are the binding constraint; for API-billed fleets, dollars are the meter and this paper doesn't speak that language. And the λe sensitivity — collapse at 0.3, good behavior at 0.4–0.5 — tells me the "sharp phase transition" is at least partly an artifact of the coefficient sweep, so I'd want to see the operating point hold across pools before trusting it. ## Overall assessment A valuable paper for the measurement discipline it brings — per-query energy profiling is the missing cost dimension in routing research — with honest failure-mode reporting and a genuinely interesting finding about RL cheapening the router itself. Recommend with revisions: evaluate under realistic serving (batching, multi-tenancy), include controller energy in the headline comparison, and test with frontier/API models in the pool where routing economics actually apply. Competing interests The author declares that they have no competing interests. Use of Artificial Intelligence (AI) The author declares that they used generative AI to come up with new ideas for their review.

View source

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Conference Sep 2010

Exploring the Sources of Waste in Kanban Software Development Projects

The application of agile software methods and more recently the integration of Lean practices contribute to the trend of continuous improvement in the software industry. One such area warranting proper empirical evidence is a project’s operational efficiency when using the Kanban method. This short paper takes a new an...

Marko Ikonen, Petri Kettunen, Nilay V. Oza et al. · 67 citations · ⚡9

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.