PREreview of "TRACER: Trace-Based Adaptive Cost-Efficient Routing for LLM Classification"
Abstract
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23179235. ## Summary This paper tackles one of the most economically significant problems in production LLM deployment: classification endpoints where every query currently pays full LLM price. The core observation is simple and correct — every production call already produces a labeled input-output pair sitting in the logs, which is a free, ever-growing training set. TRACER trains a lightweight surrogate model on these traces and governs its deployment through a "parity gate": the surrogate only serves traffic when its agreement with the LLM teacher exceeds a user-specified threshold. Interpretability artifacts describe which input regions the surrogate handles, where it plateaus, and why it defers. On a 77-class intent benchmark with a Sonnet 4.6 teacher, the system achieves 83–100% surrogate coverage depending on the threshold; on a 150-class benchmark the surrogate fully replaces the teacher. Most notably, on a natural language inference task the parity gate refuses deployment because the embedding representation cannot support reliable separation. The system is open source. ## Strengths The parity gate is the right abstraction, and the NLI refusal is the most honest result in the paper. Most cost-optimization work reports only the wins; a system that demonstrably says "not yet" when the representation is inadequate is a deployment gate worthy of the name. The interpretability artifacts address the operational question every team actually asks rather than hiding behind a single aggregate metric. Using production traces as free training data is the correct economic framing. Open-sourcing the system meaningfully increases the chance this gets adopted rather than merely cited. ## Major concerns and questions 1. Agreement with the teacher is not correctness. The parity gate measures surrogate–LLM agreement, a fidelity metric, not a quality metric. If the teacher is systematically wrong on some slice of traffic, the surrogate inherits those errors with high parity scores. The paper needs either an independent correctness check or an explicit discussion of teacher-error propagation. 2. How should a practitioner set the threshold? There is no guidance connecting it to business risk. A short section mapping the threshold to decision-layer consequences would make this deployable rather than merely publishable. 3. The 83–100% coverage range is wide and unexplained. A practitioner needs to know which side of that range their workload will land on before investing in the pipeline. 4. Distribution shift over time: when the input distribution drifts, does the gate detect degradation and re-close, or does the surrogate keep serving stale confidence? 5. All results use one teacher model. A second teacher would strengthen the generalizability claim. ## Minor points - A sentence on whether the parity-gate idea extends to generation or extraction tasks would help readers scope the contribution. - Translating coverage percentages to dollars-per-million-queries would make the economic case concrete. ## Overall assessment A strong, practitioner-relevant contribution. The parity gate plus the demonstrated willingness to refuse deployment is genuinely good engineering. Addressing teacher-error propagation and threshold-setting guidance would make this a paper I would hand to teams, not just cite. Competing interests The author declares that they have no competing interests. Use of Artificial Intelligence (AI) The author declares that they used generative AI to come up with new ideas for their review.