Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
Standard evaluation of large language models is challenged by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels, evaluating four models on three reasoning benchmarks, and finding three findings that argue for budget-conditioned evaluation protocols.