LLMs' Inner Voice Illusions and Disagreements Can Be Converted into a New Tool for Teaching Deductive Reasoning
Abstract
Contemporary education is increasingly challenged by stagnant student performance in mathematics and science, characterized by a superficial grasp of core concepts and limited critical thinking. This crisis is further intensified by the proliferation of opaque answers from artificial intelligence systems. However, the rise of Large Language Models (LLMs) utilizing Chain-of-Thought (CoT) reasoning—acting as a digital "inner voice"—offers a semi-transparent window into complex machine deliberation. We study a very specific but critical reasoning task: exclusive disjunctions between implications featuring shared consequents and opposing antecedents—a structure typical of competing scientific explanations. We address two pivotal research questions: the extent to which these LLMs reliably resolve it and whether they mirror human cognitive illusions in this task. The methodology involved evaluating six distinct LLMs across 15 task variants, while controlling for diverse contexts, to measure both conclusion accuracy and explanation correctness. Our findings indicate that while LLMs are improving, they remain inconsistent in accuracy. Notably, even when reaching correct conclusions, their argumentative paths vary, and their justifications frequently contain errors that resemble human-like logical fallacies. Nevertheless, these discrepancies may offer pedagogical opportunities. By exposing the dialogic reasoning of LLMs, teachers can facilitate structured adversarial collaborations. Students can then practice identifying cognitive biases and engage with conflicting compelling explanations, thereby fostering deductive reasoning and critical thinking.