Skip to content
Book Open access

Mitigating Neuro-Symbolic Reasoning Shortcuts with Data-Driven Knowledge Augmentation

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 2861-2872 · 0 citations · 12 references

Abstract

Recent advancements in neuro-symbolic learning (NeSy) have shown significant promise in integrating deep learning with symbolic reasoning, offering both interpretability and generalization. However, the prevalence of reasoning shortcuts, where the NeSy system predicts incorrect intermediate concepts while maintaining high final accuracy, poses a substantial challenge. This is especially problematic in domains requiring reliable and transparent decision-making. Inspired by recent theories, we find that existing methods fail to address the reasoning shortcut issue when the knowledge base lacks sufficient complexity, highlighting their vulnerability in real-world applications. In this work, we present a novel method called DKA to address this issue. It introduces a limited set of concept-supervised data to enhance the knowledge base, effectively solving the reasoning shortcut problem and improving the applicability of the NeSy system. Theoretical analysis reveals that DKA can reduce shortcut risks with improved data efficiency. Empirical studies across multiple tasks within various neuro-symbolic frameworks also verify the effectiveness of the DKA method.

Read PDF

Similar papers

Review Open access Jul 2026

Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts

This overview addresses this issue by providing a gentle introduction to RSs, discussing their causes and consequences in intuitive terms, and details methods for dealing with RSs, including mitigation and awareness strategies, and maps their benefits and limitations.

E. Marconato, Samuele Bortolotti, Emile van Krieken et al. · 1 citation
Book Open access Aug 2026

When Logic Meets Perception: Operator-Agnostic Differentiable Reasoning for Reliable Neural Prediction

Neural models in high-stakes domains lack access to ontological domain constraints that practitioners take for granted, and retrofitting such knowledge is hard: expressive logical formalisms do not scale, while scalable ones cannot express the negation, disjunction, and quantification that real constraints require. We present a differentiable reasoning framework that resolves this tension. Operating within a decidable logic that retains full Boolean expressivity, it transforms domain rules into a training objective with guaranteed polynomial-time cost per iteration. The framework is operator-agnostic - it decouples logical structure from the choice of underlying continuous semantics, revealing, through the first controlled comparison of its kind, that this choice alone can swing performance by over 30 points on the same task. This finding motivates two adaptive mechanisms: a semantic gate that focuses gradient signal on the model's most flagrant logical violations, and a structure-aware loss that automatically reweights its objective according to the logical complexity of the input constraints. Together, they eliminate the need for per-dataset loss tuning. On eight benchmark ontologies, the framework achieves statistically significant improvements over nine baselines spanning neuro-symbolic, geometric, and probabilistic paradigms. On semantic image interpretation, it refines a frozen object detector using domain rules alone - without the need for extra labels - lifting macro-averaged F1 by up to 7.8%, showing that structured knowledge, properly injected, can turn brittle pattern-matching into logically coherent prediction.

Zi-Han Shao, Chang Lu, Renate A. Schmidt et al. · 0 citations

SIRD: Symbolic Integration Rules Dataset

A novel and interpretable approach to perform symbolic integration using deep learning through integral rule prediction to speed up the search and introduces the first-of-its-kind symbolic integration rules dataset comprising two million distinct functions and integration rule pairs.

Vaibhav Sharma, M. F. Balin · 4 citations · ⚡1
#artificial intelligence Preprint Aug 2026

Stratified Consistency Distillation for Natural Language Formalization

A fine-tuning-based Stratified Consistency Distillation approach that shows significant and consistent improvements in both Pass@K and the novel Equivalent Logical Similarity metrics, demonstrating the potential of advancing logical translation through consistency distillation.

Zhi-Chao Hou, Ferhat Erata, Joseph Lilien et al. · 1 citation
Conference Jul 2026

Lightweight reasoning models for NER

Lite-CoNER is proposed, a lightweight NER framework that achieves an effective balance between recognition accuracy and inference efficiency and provides a transparent view of the decision-making process, proving that lightweight models can effectively inherit complex logic through structured distillation.

Yang Wang, Lushuang Gao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.