While large language models (LLMs) have achieved impressive gains in commonsense reasoning, they often fall into “associative shortcuts,” failing to distinguish correct answers from plausible but constraint-violating hard negatives. This reliance on semantic priors rather than specific situational constraints limits their fine-grained reasoning capabilities. To address this issue, we propose EcoReason (Evolved Commonsense Reasoning), a graph-guided evolutionary and negative-aware reinforcement learning (RL) framework. First, we introduce Graph-Guided Data Evolution, an iterative data generation strategy coupled with the student model’s training progress. In each round, we use knowledge graphs (KGs) to identify deceptive sibling concepts and employ a teacher LLM to create constraint-heavy question answering (QA) data targeting the student’s current blind spots. As training progresses, the synthesized curriculum becomes increasingly challenging. Second, we propose Negative-Aware Policy Optimization (NAPO), an RL algorithm built upon Group Relative Policy Optimization (GRPO). NAPO identifies “stubborn negatives,” defined as incorrect options that are repeatedly selected by sampled policies within the same rollout group, and applies stronger penalties to these recurring distractor-specific errors. Experiments show that EcoReason substantially improves LLM commonsense reasoning in complex constraint-heavy scenarios.
Xin Guan, Jiuxin Cao, Biwei Cao et al.· IEEE Transactions on Audio,...· 0 citations
SARK is proposed, a style-augmented multi-task framework that prioritizes effective knowledge over stylistic perturbations in the reranker model and improves generation performance across multiple LLMs under mixed-style conditions.
Ruwen Zhang, Bo Liu, Zhang-Sheng Xiang et al.· Annual Meeting of the Associ...· 0 citations
The LLM-Enhanced Component Dependency Evolution Graph (CDEG) framework is proposed, a hybrid representation that fuses structural features extracted by Tree-sitter with semantic embeddings derived from a fine-tuned LLM, effectively distinguishing backported patches from code refactoring.
Yuan-Jun Gao, Hong-Zhou Wu, Yu-Jia Luo et al.· Mathematics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.