Skip to content
Conference Open access

SMART: Evaluating LLMs' Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark

2026 · Annual Meeting of the Association for Computational Linguistics · pp. 35426-35452 · 0 citations · 39 references
Computer Science

TL;DR

This work proposes SMART, a benchmark that decomposes mathematical problem-solving into four cognitive dimensions: S emantic Understanding, M athematical Reasoning, A rithmetic Computation, and R eflection & Refinemen t, and introduces dimension-specific tasks to measure the corresponding cognitive processes of LLMs.

Abstract

Large Language Models (LLMs) have achieved remarkable performance across a wide range of mathematical benchmarks. However, concerns remain as to whether these successes reflect genuine reasoning or superficial pattern recognition. Existing evaluation methods, which typically focus either on the final answer or on the intermediate reasoning steps, reduce mathematical reasoning to a shallow input–output mapping, overlooking its inherently multi-stage and multi-dimensional cognitive nature. Inspired by Pólya’s problem-solving theory, we propose SMART, a benchmark that decomposes mathematical problem-solving into four cognitive dimensions: S emantic Understanding, M athematical Reasoning, A rithmetic Computation, and R eflection & Refinemen t , and introduces dimension-specific tasks to measure the corresponding cognitive processes of LLMs. We apply SMART to 22 state-of-the-art open-and closed-source LLMs and uncover substantial discrepancies in their capabilities across dimensions. Our findings reveal genuine weaknesses in current models and motivate a new metric, the All-Pass Score, designed to better capture true problem-solving capability. Data is available at https://huggingface.co/datasets/ewdfd/SMART.

Read PDF

Similar papers

2025

Evaluating the Inductive Abilities of Large Language Models: Why Chain-of-Thought Reasoning Sometimes Hurts More Than Helps

This work presents a theoretical framework that reveals how reasoning steps can amplify error through three failure modes: incorrect sub-task decomposition, incorrect sub-task solving, and incorrect final answer summarization, and introduces structured interventions that adapt CoT generation according to the identified failure types.

Haibo Jin, Peiyan Zhang, Man Luo et al. · 1 citation
Preprint Aug 2026

Cognitive Profiling of LRMs'Reasoning Traces Using Bloom's Taxonomy

This work introduces a framework for automatic annotation of reasoning steps through the lens of Bloom's Taxonomy, which classifies thinking into six cognitive levels, such as Remembering, Applying and Evaluating, and demonstrates that thinking-type information derived from reasoning traces correlates with correctness, paving the way for improved reasoning.

Maria-Eleni Zoumpoulidi, Georgios Paraskevopoulos, A. Potamianos · 0 citations
Preprint Aug 2026

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs'Agentic Mathematical Capabilities

Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles, demonstrating that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.

Jiayi Kuang, Yinghui Li, Yun-Ze Song et al. · 0 citations
#natural language process... Preprint Aug 2026

INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

INSPIRE is an Internalize-Then-Improve approach combining Reference-Guided Student Internalization (RGSI), which produces high-quality preference candidates under the policy model's own distribution, with a stage-wise rubric preference training strategy that decomposes learning into method-oriented and correctness-oriented stages.

Shuai Wang, Jiayi Kuang, Yinghui Li et al. · 0 citations
Review Open access Jul 2026

A Survey of Commonsense Reasoning in LLMs

Commonsense reasoning plays a crucial role in natural language processing (NLP) by enabling systems to determine causation and assess the plausibility of various scenarios. Despite significant research focused on formalizing commonsense knowledge and integrating it with mathematical logic, replicating such reasoning in artificial intelligence systems—such as large language models (LLMs)—remains a formidable challenge. Despite these challenges, LLMs demonstrate significant promise and exhibit ongoing progress in reasoning capabilities. This article surveys the current state of commonsense reasoning in LLMs, including datasets, models, benchmarks, enhancements, opportunities, and challenges—highlighting both advancements and limitations. We discuss the role of commonsense reasoning in LLMs, how they strive to emulate human-like reasoning, and why they often struggle to capture the nuanced inferences associated with commonsense understanding of the physical and social world. This survey emphasizes the need for continued investigation and optimization, including the integration of mechanisms to enhance the reliability of these models. Overall, while LLMs are making strides in commonsense reasoning, overcoming their limitations requires further research and development to fully harness their potential.

N. Teo, Donghao Huang, Zhaoxia Wang · 0 citations
Preprint Aug 2026

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential branching.

Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.