Skip to content

Multi-model relevance-guided reverse reasoning: improving chain-of-thought with small models

Aug 2026 · Journal of Supercomputing · Vol 82 · 0 citations · 58 references

TL;DR

This work proposes Multi-model Relevance-Guided Reverse Reasoning (MRGRR), which uses a target model and auxiliary models to generate reverse reasoning prompt candidates and shows higher average accuracy than six CoT baselines.

View source

Similar papers

Preprint Aug 2026

ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning

This work proposes ChainPrune, a novel reasoning path semantic structural optimization method to efficiently and controllably synthesize self-generated high-quality training data and incorporates a DPO-based preference learning method combined with supervised loss, effectively mitigating false reward suppression.

Weihang Pan, Zhengxu Yu, Yuxiang Zhang et al. · 1 citation
2025

Evaluating the Inductive Abilities of Large Language Models: Why Chain-of-Thought Reasoning Sometimes Hurts More Than Helps

This work presents a theoretical framework that reveals how reasoning steps can amplify error through three failure modes: incorrect sub-task decomposition, incorrect sub-task solving, and incorrect final answer summarization, and introduces structured interventions that adapt CoT generation according to the identified failure types.

Haibo Jin, Peiyan Zhang, Man Luo et al. · 1 citation
Book Open access Jul 2026

Think, But Don't Tell: Implicit Reasoning for LLM-based Sequential Recommendation via Multi-Teacher Distillation

I Reasoning via Multi-Teacher Distillation is proposed, a novel framework that 'compiles' the reasoning abilities of large teacher LLMs into a lightweight student Small Language Model (SLM), which significantly outperforms state-of-the-art baselines in both recommendation accuracy and inference efficiency.

Weihai Lu, Xiaoxi Cui, Chenke Yin · 0 citations
Review Open access 2026

Negative Contrastive Chain-of-Thought Distillation for Transfer Reasoning Capability to Small Language Models

Large language models (LLMs) demonstrate strong reasoning capabilities through chain-of-thought prompting. Recently, numerous studies have focused on transferring reasoning capability from LLMs to small language models (SLMs) through chain-of-thought (CoT) distillation. However, these approaches face two limitations. First, simplistic multiple iterations to augment correct predictions are ineffective, due to the LLM tends to reproduce the same errors. Second, directly utilizing incorrect predictions during the training process exposes the model to learning incorrect predictions of the LLM during distillation. To address these problems, we propose a novel framework named negative contrastive CoT distillation (NCD), which efficiently augments correct predictions while training an SLM to exclude incorrect predictions. NCD comprises of two stages: corrective reviewer augmentation (CRA) and negative exclusive learning (NEL). CRA utilizes the self-correction capability of LLM to augment the correct predictions within a single iteration. NEL is a distillation strategy that strengthens the ability of SLM to mimic the correct predictions of LLM by leveraging the incorrect predictions of LLM. The experimental results show that NCD improves the reasoning capability of the SLM by a sizable gap in diverse reasoning tasks, including mathematical and commonsense reasoning tasks. We demonstrate both the efficiency of CRA and the effectiveness of NEL.

Jae-Wook Han, Hayoung Jo, Jung-Ho Hong et al. · 0 citations
Jul 2026

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning

REFACT is an adaptive fact-restatement citation framework that enables LLMs to determine when contextual grounding is needed and selectively restate source facts at appropriate levels of detail for reliable reasoning.

Zhensheng Jin, Xin Dai, Zhenghao Liu et al. · 0 citations
#small language model Open access Aug 2026

Are reasoning paradigms scale-aware? A cross-paradigm verification of prompting, retrieval, and knowledge-graph scaffolding for small language models

A scale-aware comparative study of reasoning enhancement for SLMs across three major families of methods: prompting-based reasoning, retrieval-based augmentation, and knowledge graph guided scaffolding shows that reasoning-enhancement strategies are not universally transferable across model scales under the evaluated settings.

Zhen-Zhen Gu, Jie Liu, Xian Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.