Reasoning Budget Sweep: truncation and answer-production measurements for 12 small reasoning LLMs
A per-response measurement dataset covering 12 open reasoning language models (0.6B to 7B) on AIME 2024, AIME 2025 and MATH-500 across six generation token budgets (512 to 16384) and three seeds. 11,118 scored responses. Each row records the token count, whether the response was truncated at the budget, whether an answer could be extracted, and whether it was correct. Intended for meta-analysis of how generation budgets interact with benchmark scores in the small-model regime. Chain-of-thought text is not included.