Aug 2026· Journal of Supercomputing· Vol 82· 0 citations· 58 references
TL;DR
This work proposes Multi-model Relevance-Guided Reverse Reasoning (MRGRR), which uses a target model and auxiliary models to generate reverse reasoning prompt candidates and shows higher average accuracy than six CoT baselines.
This work proposes ChainPrune, a novel reasoning path semantic structural optimization method to efficiently and controllably synthesize self-generated high-quality training data and incorporates a DPO-based preference learning method combined with supervised loss, effectively mitigating false reward suppression.
Weihang Pan, Zhengxu Yu, Yuxiang Zhang et al.· 1 citation
This work presents a theoretical framework that reveals how reasoning steps can amplify error through three failure modes: incorrect sub-task decomposition, incorrect sub-task solving, and incorrect final answer summarization, and introduces structured interventions that adapt CoT generation according to the identified failure types.
Haibo Jin, Peiyan Zhang, Man Luo et al.· Neural Information Processin...· 1 citation
I Reasoning via Multi-Teacher Distillation is proposed, a novel framework that 'compiles' the reasoning abilities of large teacher LLMs into a lightweight student Small Language Model (SLM), which significantly outperforms state-of-the-art baselines in both recommendation accuracy and inference efficiency.
Large language models (LLMs) demonstrate strong reasoning capabilities through chain-of-thought prompting. Recently, numerous studies have focused on transferring reasoning capability from LLMs to small language models (SLMs) through chain-of-thought (CoT) distillation. However, these approaches face two limitations. First, simplistic multiple iterations to augment correct predictions are ineffective, due to the LLM tends to reproduce the same errors. Second, directly utilizing incorrect predictions during the training process exposes the model to learning incorrect predictions of the LLM during distillation. To address these problems, we propose a novel framework named negative contrastive CoT distillation (NCD), which efficiently augments correct predictions while training an SLM to exclude incorrect predictions. NCD comprises of two stages: corrective reviewer augmentation (CRA) and negative exclusive learning (NEL). CRA utilizes the self-correction capability of LLM to augment the correct predictions within a single iteration. NEL is a distillation strategy that strengthens the ability of SLM to mimic the correct predictions of LLM by leveraging the incorrect predictions of LLM. The experimental results show that NCD improves the reasoning capability of the SLM by a sizable gap in diverse reasoning tasks, including mathematical and commonsense reasoning tasks. We demonstrate both the efficiency of CRA and the effectiveness of NEL.
Jae-Wook Han, Hayoung Jo, Jung-Ho Hong et al.· IEEE Access· 0 citations
REFACT is an adaptive fact-restatement citation framework that enables LLMs to determine when contextual grounding is needed and selectively restate source facts at appropriate levels of detail for reliable reasoning.
Zhensheng Jin, Xin Dai, Zhenghao Liu et al.· arXiv.org· 0 citations
A scale-aware comparative study of reasoning enhancement for SLMs across three major families of methods: prompting-based reasoning, retrieval-based augmentation, and knowledge graph guided scaffolding shows that reasoning-enhancement strategies are not universally transferable across model scales under the evaluated settings.
Zhen-Zhen Gu, Jie Liu, Xian Liu· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.