Skip to content
Conference Open access

Thesis Proposal: Intentional Inference for Insight Generation

2026 · Annual Meeting of the Association for Computational Linguistics · pp. 1236-1252 · 0 citations · 112 references
Computer Science

TL;DR

This work examines how models can identify missing premises and surface multiple plausible interpretations to make evaluation more rigorous, and explores how to improve reasoning to enable deeper inferences, focusing on code generation and qualitative reasoning.

Abstract

Large language models (LLMs) show strong capabilities in natural language generation (NLG) and have been applied to translate complex structured data into human-readable insights. While these models excel at surface-level fluency, they remain unreliable as they produce factually inaccurate outputs and struggle with consistent logical inference beyond surface-level patterns. Moreover, they often lack a clear sense of relevance and produce shallow or un-informative insights. This proposal argues that a key source of these limitations is task underspecification, which requires models to make implicit assumptions about missing context. We investigate how such underspecification leads to unintentional assumptions and how these affect faithfulness and evaluation. We examine how models can identify missing premises and surface multiple plausible interpretations to make evaluation more rigorous. We also explore how to improve reasoning to enable deeper inferences, focusing on code generation and qualitative reasoning. Finally, we will evaluate how the underlying assumptions and depth of inference influence the perceived interestingness of the insights. By shifting focus from surface-level generation to assumption-aware deeper inferences, this work aims to improve reliability, interpretability, and user controllability in NLG.

Read PDF

Similar papers

Book Open access Jul 2026

Attend to Fragments: How Key Information Affects Large Language Models for Factual Inconsistency Detection

A new benchmark, KIFI, is designed, which comprises 1032 carefully selected instances from the TRUE and ScreenEval datasets, with key information annotated, and it is shown that LLMs frequently fail to use the appropriate information to make correct decisions.

Xindi Guo, Zhen Xie, Patrick H. Chen · 0 citations
#natural language process... Preprint Sep 2026

Do LLMs Make More Mistakes If They Do Not Believe the Input Data?

Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context depends on how plausible they perceive the context to be (context-memory conflict). To better identify error patterns, we make use of the increased difficulty of non-English and low-resource language text generation and input data based on local knowledge, only partially captured in models'parametric knowledge. We let the models generate text in English, Czech, Slovak and Upper Sorbian from factual (FA), counterfactual (CFA) and fictional (FI) RDF triples containing local Czech and Slovak data. Contrary to our expectations, we observe only a weak context-memory conflict on the human-annotated sample. For Kimi K3 as an LLM judge, which agrees well with human annotations on the sample, counterfactual inputs receive only slightly lower faithfulness scores than factual ones (-0.05 on a 1-5 scale). We also find that a suboptimal choice of LLM judge would lead to overestimating the strength of the context-memory conflict.

Peter Kochelka, Ale\v{s} Manuel Pap\'a\v{c}ek, Vojt\v{e}ch Dvo\v{r}\'ak et al. · 0 citations
Book Open access Jul 2026

Tokens to Types: Context Editing with Selective Entity Abstraction for Grounded Generation

This framework proposes a context-editing framework that performs selective abstraction over entities that appear in both the context and the question, establishing symbolic abstraction as a highly cost-efficient solution for ensuring context fidelity in LLMs.

Rounak Sharma, Debabrata Mahapatra, S. Saini · 0 citations
Open access Aug 2026

Towards Trustworthy Large Language Models

An integrated conceptual frame-work that couples attention- and perturbation-based explainability with lightweight hallucination-detection signals and token-efficient inference strategies is presented, and a set of cross-cutting consistency metrics are instrumented with a set of cross-cutting consistency metrics.

Sakshi Parate, Shreyans Sanyal · 0 citations
Jul 2026

Implicit Reasoning Steering via Concept Chaining

The results show that indirect, natural-looking text can systematically steer model predictions while remaining substantially less inferable than direct paraphrases, which shows that reasoning brittleness is not merely an evaluation artifact: it creates a practical channel through which latent biases can be amplified by ordinary-looking text to covertly redirect model decisions.

Xiao Ye, Sanika Chavan, Yuxi Huang et al. · 0 citations
#machine learning Preprint Sep 2026

Evaluation of Contextual Understanding in Large Language Models

Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs extract, integrate, and reason over contextual information--a gap particularly critical in question answering, where models must align responses with contextually grounded knowledge rather than memorized associations. We propose a novel knowledge graph-based evaluation framework introducing Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure integrating structural and semantic similarity into a continuous evaluation score, alongside a diagnostic framework for categorizing reasoning errors. To validate this pipeline, we evaluate S3KG against established metrics on a curated question-answer (QA) benchmark, demonstrating its effectiveness in measuring correctness, faithfulness, and interpretability in LLM-generated responses.

Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.