Skip to content
Review

Conversational Grounding in Large Language Models: Evaluation Methods, Challenges and Future Directions

2026 · SIGDIAL Conferences · pp. 151-163 · 0 citations · 66 references
Computer Science

TL;DR

This paper surveys how conversational grounding is evaluated in task-oriented dialogue in the current era of LLMs and focuses on how conversational grounding is modelled explicitly—using dialogue acts and by modelling the participant mental state.

View source

Similar papers

Book Open access Sep 2026

Conversational Style in Open Domain Dialogue Systems: What Makes a Response Sound Natural

Intelligent Virtual Agents need to be able to participate in extended dialogue interactions while maintaining a conversational style. We discover the elements of conversational style in open-domain dialogues by analyzing the features that distinguish conversational system responses from responses that are merely topically relevant. We first collect 7,751 dialogue contexts from live human conversations with a multi-generator Alexa Prize SocialBot that was deployed across four competition years. We also collect 35,623 candidate responses for the contexts. We annotate the responses with a four-level ABCD quality scheme that isolates conversational naturalness from topical relevance. We then extract twenty-five linguistic features that capture conversational properties of responses and contrast (A) responses that have a conversational style, from (B) responses that are topically relevant, but less conversational. We find that conversational responses are marked by other-directed engagement: question-asking, user engagement phrases, acknowledgment openings, and second-person reference, while responses that are merely relevant and topical are marked by self-oriented information delivery: opinion markers, formulaic openings, and hedges and emphasizers deployed in service of the system’s own assertions. We thus find that conversational style in this setting is best understood as a pragmatic orientation toward the user rather than toward the system’s own content, and that the system must be mixed-initiative to manifest a conversational style. We discuss what these findings imply for the design of Intelligent Virtual Agents.

Vrindavan Harrison, M. Walker · 0 citations
Book Open access Aug 2026

Post-Thinking in NPC Dialogue: A Paradigm for Reflective Character Models

Immersive and believable NPC dialogue requires characters that feel intentional. They remember specific information about themselves, follow through on their goals, and stay true to their personalities across long conversations. We introduce Post-Thinking, a technique that maintains a rolling reflection trace across chat turns. After each response, effectively in the dead time between LLM queries, the model generates a trace reflecting its current goals, emotional state, and narrative intentions. This trace is kept in context and conditions the next response, directing the conversation while designed to add zero perceivable latency to the end user. Critically, each trace is generated with the previous n reflection traces still visible in context, allowing the character’s inner state to compound and evolve naturally throughout extended conversations. As a preliminary study, we synthetically annotated conversations from seed datasets and interviewed human experts to assess their quality in preparation for fine-tuning. We expect the reflection generation pass to help actively ground the character as it encourages the model to explicitly surface aspects of the character definition most relevant to the current moment, counteracting the prompt drift that typically degrades consistency over long exchanges.

Keegan Carey, Hexi Wang · 0 citations
#natural language process... Preprint Aug 2026

You Know What I Mean: A Benchmark for Agentic Conversational Reference Grounding

The results show that CoRG remains challenging for current agents, even the best agent reaches only 67.0% success rate, leaving one third of references unresolved, and position CoRG as a concrete benchmark for studying how agents search, inspect, and verify information in realistic multi-tool environments.

Karen Fuchs, Uri Katz, Yoav Goldberg · 0 citations
Review Open access Aug 2026

Benchmark Research on Safety Value Alignment Evaluation of Open-Domain Dialogue Systems based on NLP

Large language models sometimes behave in puzzling ways. They pass various safety tests, yet in multi-turn dialogues, a few carefully crafted sentences can lead them astray into making dangerous judgments. We call this "extreme value alignment failure." A review of recent research reveals an awkward situation: attack, defense, evaluation, and theoretical studies operate in isolation, with little connection among them. This fragmentation results in repeated extreme risks that remain unresolved. This paper maps the four research directions onto a unified framework---"failure mode, attack vector, defense level, evaluation benchmark"---providing a theoretical coordinate for the field and directions for future evaluation research.

Pei-Rong Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.