An exam-style evaluation framework is introduced for studying the global budget allocation of reasoning language models when multiple problems share an end-to-end cost or latency constraint, in which a model must distribute one shared token budget across questions with different difficulty and point values.
Abstract
Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency constraint, models must decide how to divide limited inference compute among them. We introduce an exam-style evaluation framework for studying this setting, in which a model must distribute one shared token budget across questions with different difficulty and point values to maximize its total score. Across several open and frontier reasoning models, we find that models fail to allocate a shared budget strategically across questions of varying difficulties and values. Models behave largely as greedy sequential solvers: they prioritize questions by presentation order, front-load effort on early questions, and remain insensitive to value, with these tendencies becoming more pronounced as the number of questions grows. Explicit planning prompts spread compute more evenly but do not produce value- or difficulty-aware prioritization. The same behavioral pattern extends from mathematical to code reasoning. These findings establish global budget allocation as a distinct capability that is not captured by conventional per-question evaluation and remains a challenge for current reasoning models.
By periodically pruning unpromising trajectories and immediately branching from high-quality prefixes, Gambit dynamically concentrates compute onto the most promising reasoning traces via a light-weight scorer probing hidden states while maintaining continuous high hardware utilization.
Lijie Yang, Hongyin Luo, Tri Dao et al.· 0 citations
The first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing is conducted - grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation.
Davide Romano, Kanak Raj, Jerrod Parker et al.· 0 citations
Analysis shows that BRANCH's advantage arises not only from exploring multiple reasoning paths, but also from recovering from truncation: its gains strongly correlate with the baseline rate of empty, budget-exhausted outputs, weakening the hypothesis that different problems require routing among test-time reasoning operators.
Sheng Zhang, Xiao-Min Wu, Xiyang Wu et al.· 0 citations
This paper investigates whether there exists a token-budget threshold, below which the overhead of planning and verification hurts performance and above which it helps, and evaluates two systems on FinQA and TAT-QA financial reasoning tasks.
Thomas Nolasque, John Grey, Calista Pham et al.· 0 citations
This work presents a theoretical framework that reveals how reasoning steps can amplify error through three failure modes: incorrect sub-task decomposition, incorrect sub-task solving, and incorrect final answer summarization, and introduces structured interventions that adapt CoT generation according to the identified failure types.
Haibo Jin, Peiyan Zhang, Man Luo et al.· Neural Information Processin...· 1 citation
Paired runs of the same question at matched reasoning horizons across 198 GPQA Diamond and 500 MMLU-Pro questions are evaluated, finding that a tight deadline can favor lower effort or a concise instruction, whereas allowing high effort to finish can recover higher final accuracy.
Francesca Carlon, Vincent Ginis, A. Algaba· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.