2026· Annual Meeting of the Association for Computational Linguistics· pp. 35426-35452· 0 citations· 39 references
Computer Science
TL;DR
This work proposes SMART, a benchmark that decomposes mathematical problem-solving into four cognitive dimensions: S emantic Understanding, M athematical Reasoning, A rithmetic Computation, and R eflection & Refinemen t, and introduces dimension-specific tasks to measure the corresponding cognitive processes of LLMs.
Abstract
Large Language Models (LLMs) have achieved remarkable performance across a wide range of mathematical benchmarks. However, concerns remain as to whether these successes reflect genuine reasoning or superficial pattern recognition. Existing evaluation methods, which typically focus either on the final answer or on the intermediate reasoning steps, reduce mathematical reasoning to a shallow input–output mapping, overlooking its inherently multi-stage and multi-dimensional cognitive nature. Inspired by Pólya’s problem-solving theory, we propose SMART, a benchmark that decomposes mathematical problem-solving into four cognitive dimensions: S emantic Understanding, M athematical Reasoning, A rithmetic Computation, and R eflection & Refinemen t , and introduces dimension-specific tasks to measure the corresponding cognitive processes of LLMs. We apply SMART to 22 state-of-the-art open-and closed-source LLMs and uncover substantial discrepancies in their capabilities across dimensions. Our findings reveal genuine weaknesses in current models and motivate a new metric, the All-Pass Score, designed to better capture true problem-solving capability. Data is available at https://huggingface.co/datasets/ewdfd/SMART.
This work presents a theoretical framework that reveals how reasoning steps can amplify error through three failure modes: incorrect sub-task decomposition, incorrect sub-task solving, and incorrect final answer summarization, and introduces structured interventions that adapt CoT generation according to the identified failure types.
Haibo Jin, Peiyan Zhang, Man Luo et al.· Neural Information Processin...· 1 citation
This work introduces a framework for automatic annotation of reasoning steps through the lens of Bloom's Taxonomy, which classifies thinking into six cognitive levels, such as Remembering, Applying and Evaluating, and demonstrates that thinking-type information derived from reasoning traces correlates with correctness, paving the way for improved reasoning.
Maria-Eleni Zoumpoulidi, Georgios Paraskevopoulos, A. Potamianos· 0 citations
Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles, demonstrating that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.
Jiayi Kuang, Yinghui Li, Yun-Ze Song et al.· 0 citations
INSPIRE is an Internalize-Then-Improve approach combining Reference-Guided Student Internalization (RGSI), which produces high-quality preference candidates under the policy model's own distribution, with a stage-wise rubric preference training strategy that decomposes learning into method-oriented and correctness-oriented stages.
Shuai Wang, Jiayi Kuang, Yinghui Li et al.· 0 citations
Commonsense reasoning plays a crucial role in natural language processing (NLP) by enabling systems to determine causation and assess the plausibility of various scenarios. Despite significant research focused on formalizing commonsense knowledge and integrating it with mathematical logic, replicating such reasoning in artificial intelligence systems—such as large language models (LLMs)—remains a formidable challenge. Despite these challenges, LLMs demonstrate significant promise and exhibit ongoing progress in reasoning capabilities. This article surveys the current state of commonsense reasoning in LLMs, including datasets, models, benchmarks, enhancements, opportunities, and challenges—highlighting both advancements and limitations. We discuss the role of commonsense reasoning in LLMs, how they strive to emulate human-like reasoning, and why they often struggle to capture the nuanced inferences associated with commonsense understanding of the physical and social world. This survey emphasizes the need for continued investigation and optimization, including the integration of mechanisms to enhance the reliability of these models. Overall, while LLMs are making strides in commonsense reasoning, overcoming their limitations requires further research and development to fully harness their potential.
The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential branching.
Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.