PREreview of "Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving"
Abstract
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23177620. ## Summary Routing papers usually optimize proxy costs — parameter counts, API prices, model counts. This one measures the thing itself: per-query GPU energy. The authors run an offline tournament over a fixed pool of small-to-mid open models, recording correctness, latency, board power (polled via pynvml at 200ms, idle baseline subtracted), and effective GPU energy for each query. A small LLM router is then trained on these measurements — supervised fine-tuning followed by GRPO with an energy coefficient λe in the reward — to select one answer model per query. On 913 held-out questions across seven benchmarks, the KL-GRPO checkpoint matches UniRoute (60.9% vs 61.8% accuracy) at ~23% lower mean answer energy (1.69 vs 2.19 kJ); a GRPO checkpoint sits near Smoothie (58.4% vs 59.9%) at ~24% lower energy. The paper also reports a sharp accuracy–energy phase transition across routers, and a striking training effect: RL cuts the controller's own decode from ~2.4s to ~0.8s, making the router itself cheaper. Training cost is quantified as a one-time 0.32–0.40 kWh. ## Strengths 1. **Measured joules, not proxies — the right measurement discipline.** Energy is the cost dimension routing research has mostly ignored, and per-query pynvml profiling with idle-baseline subtraction is how you do it honestly. The finding that MATH-500 is 13% of questions but 52–57% of energy is exactly the kind of measurement that changes how a practitioner thinks about their workload mix. 2. **Admirably honest about failure modes.** The paper reports policy collapse outright — GRPO at λe=0.3 routes literally everything to Llama 3.1 8B with zero selection entropy — plus SFT over-selecting the 0.5B model, and the mid-size-arm aggregate illusion (Always-Qwen2.5-7B looks strong at 64.0% overall but collapses to 26.2% on BBH). These are the results that make the rest of the paper believable. 3. **RL made the router itself cheaper — a genuinely interesting effect.** The 2.4s→0.8s controller decode improvement means RL didn't just pick cheaper answers, it learned to decide faster. Router overhead is the perennial objection to learned routing, and this is the first paper I've seen that measures the router's own cost shrinking under training. ## Major concerns 1. **Offline lookup on one A100 is not serving.** All evaluation is offline replay of pre-logged tournament outcomes — no live serving, no batching, no continuous batching, no multi-tenancy, one A100-SXM4-80GB via Ollama. Production energy per query depends on batching, PUE, idle-power amortization, and hardware mix; joules measured in a lab rig don't translate to anyone's bill without those. The paper's own future work admits online and multi-turn agent settings are unevaluated — that's where the energy actually lives. 2. **The controller's energy is excluded from the headline comparison.** The authors state they "do not invent a controller cost" for baselines — fair — but the main tables then report answer-model energy only, while their own Figure 5 shows controller energy in the hundreds of joules. On cheap queries the router can cost nearly as much as the answer it selects; the headline 23% saving is computed on a ledger that omits the router. 3. **The pool excludes the models where routing actually saves money.** Everything is 0.5B–32B open models; no frontier or API-billed models, and the accuracy band tops out around 60–62% while always-largest hits 73.5%. The router never reaches just-use-the-big-model accuracy, so the "savings" are measured inside a pool nobody routes between in production — the real routing decision (small vs. frontier API) is outside the experiment entirely. ## Practitioner perspective I'd buy the measurement methodology before the routing policy: run the offline tournament on your own fleet, in joules, and you'll learn more about your workload than any router paper will tell you. Energy-aware routing matters most for on-prem and GPU-constrained serving, where joules are the binding constraint; for API-billed fleets, dollars are the meter and this paper doesn't speak that language. And the λe sensitivity — collapse at 0.3, good behavior at 0.4–0.5 — tells me the "sharp phase transition" is at least partly an artifact of the coefficient sweep, so I'd want to see the operating point hold across pools before trusting it. ## Overall assessment A valuable paper for the measurement discipline it brings — per-query energy profiling is the missing cost dimension in routing research — with honest failure-mode reporting and a genuinely interesting finding about RL cheapening the router itself. Recommend with revisions: evaluate under realistic serving (batching, multi-tenancy), include controller energy in the headline comparison, and test with frontier/API models in the pool where routing economics actually apply. Competing interests The author declares that they have no competing interests. Use of Artificial Intelligence (AI) The author declares that they used generative AI to come up with new ideas for their review.