Skip to content

The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA

Sep 2026 · 0 citations · 19 references
Computer Science

TL;DR

This work isolates SLM graph reasoning from navigation and traceability capability by employing the THESEUS navigation and traceability framework and using frozen, off-the-shelf SLMs as local action policies, and results motivate evaluating SLM graph reasoning beyond endpoint accuracy alone.

Abstract

Small language models (SLMs) are increasingly paired with knowledge graphs (KGs), yet end-to-end KG question answering conflates graph access, search, navigation, reasoning, and answer generation. This coupling makes it difficult both to determine whether an SLM can faithfully execute the reasoning path implied by a question and to attribute failures to navigation rather than to other stages of the pipeline. We isolate this capability by employing the THESEUS navigation and traceability framework and using frozen, off-the-shelf SLMs as local action policies. At each hop, the environment exposes the legal outgoing graph actions, and the model selects one executable graph action and decides whether to stop, without task-specific parameter updates, model-controlled beam search, or free-form answer generation. This controlled setting allows us to evaluate terminal-answer accuracy with Hits@1 together with path fidelity, using Path Edit Distance (PED) as the primary trajectory metric. Across the Kinship and MQuAKE-ST KGQAs, similarly sized local models differ substantially in answer accuracy and path fidelity, with the two metrics sometimes favoring different models. This model-dependent behavior also extends to prompting, as a single demonstrated trajectory can improve or degrade navigation depending on the model. These results motivate evaluating SLM graph reasoning beyond endpoint accuracy alone.

View source

Similar papers

#machine learning Preprint Sep 2026

Theseus in the Graph: Towards Traceable Multi-Hop Graph Navigation

Multi-Hop Knowledge Graph Question Answering (KGQA) tasks require models to assemble relational evidence along paths in a KG to answer natural-language questions. However, existing KGQA systems typically focus on predicting the final answer without explicitly modeling or validating the intermediate reasoning steps, obs...

Eduin E. Hernandez, Luis F. Garcia, Nurassyl Askar et al. · 1 citation · ⚡1
Conference Aug 2026

Introspective Dynamic Strategy for Knowledge Graph Enhanced Reasoning in Large Language Models

Large Language Models (LLMs) achieve strong performance on many tasks, yet their reasoning can remain brittle and their decision process is often insufficiently transparent. Knowledge graphs (KGs) provide explicit entities and relations that can ground reasoning and improve traceability. However, most KG-augmented appr...

Jia-Jun Duan, Qiu-Ling Chen, Ze-Xu Wang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering

Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on knowledge-intensive tasks. Knowledge graphs (KGs) provide LLMs with structured, interpretable, and updatable factual grounding, making them a promising external knowledge sou...

Yang Hong, Ya-Jun Yang, Xin Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

This work compares country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct across Qwen, Llama, and Gemma to separate early readability, natural strength, causal steering, and later content dependence.

Wen-Lin Wei, Yuan Fang, Ren-He Jiang et al. · 0 citations
#natural language process... Preprint Oct 2026

Wikidata Search Traces: A Dataset for Training Knowledge Graph Search Agents

Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right entities and properties and chains their relations. Language models offer a natural-language alternative but answer largely from memory, which is least reliable for less promi...

Mohamed Chenene, C. Rosas Hinostroza, A. Stasenko et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.