Skip to content
Preprint

Beyond the Final Prompt: Measuring the Effect of Within-Conversation Context on AI Answers

Aug 2026 · 1 citation · 13 references
Computer Science

TL;DR

This study directly test whether that omitted within-conversation context changes answers in a conversation and concerns preceding turns in the same conversation and does not test persistent memory across separate conversations.

Abstract

An isolated final user message is often treated as the query in evaluations of AI systems. In a conversation, however, the actionable request may be distributed across preceding turns. We directly test whether that omitted within-conversation context changes answers. For each of 180 English multi-turn conversations sampled from a governed commercial corpus and the public PRISM dataset, we hold the final user message and requested answer model constant while generating three answers: one from the full role-labelled conversation, one from the final message alone, and one from the final message plus a prefix-only reconstruction capped at 160 words. A separately requested judge model evaluates answers under randomized labels. The prespecified primary endpoint is a material difference that could change what the user does, rather than a difference in style or detail. After inverse-probability weighting to the eligible cohorts, the full-conversation and isolated-final answers differ materially in 44.7% of cases (95% bootstrap CI 33.8% to 56.1%). Full-conversation answers score 0.49 points higher on a 0 to 4 request-satisfaction scale (0.32 to 0.67). Adding the compressed prefix reduces the material-difference rate to 30.8% (20.2% to 42.1%), a 13.9-point reduction (4.9% to 24.1%), and reduces the mean satisfaction gap to 0.01 points (-0.12 to 0.13). Yet compression is not equivalent to the complete dialogue context: almost one third of answers remain materially different. An order-swapped repeat on 48 cases yields 91.7% agreement and kappa = 0.83 for the primary decision. The study concerns preceding turns in the same conversation and does not test persistent memory across separate conversations.

View source

Similar papers

Jul 2026

The Prompt Is Not the Query: How Request State Evolves Across Multi-Turn AI Conversations

Categorical results support session-level measurement for AI search, and length-matched nulls show that low lexical coverage is largely a consequence of turn length, so vocabulary results are interpreted as information availability, not semantic drift.

Benjamin Tannenbaum · 2 citations
Preprint Aug 2026

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in a 4-month deployment (40 users, 1,872 sessions, 7 memory conditions). Existing-benchmark Direct QA varies from 19.7% to 70.1% across the 7 conditions, but satisfaction does not change. We hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response). To examine this, we introduce MemUse, a set of real user-cued memory moments drawn from the deployment, scored by an integration-aware judgment of the natural conversational response. Holding the model and context fixed, the same system that scores 78.8% on Direct QA references only 7.9% of those facts in conversation -- a 71-point gap. Within these moments, Natural Integration is associated with satisfaction, whereas Direct QA is not. We release the deployment corpus and MemUse together with all judgments and scoring prompts at https://github.com/ryuichi-sumida/memuse.

Ryuichi Sumida, K. Inoue, Tatsuya Kawahara · 0 citations
Book Open access Sep 2026

Conversational Style in Open Domain Dialogue Systems: What Makes a Response Sound Natural

Intelligent Virtual Agents need to be able to participate in extended dialogue interactions while maintaining a conversational style. We discover the elements of conversational style in open-domain dialogues by analyzing the features that distinguish conversational system responses from responses that are merely topically relevant. We first collect 7,751 dialogue contexts from live human conversations with a multi-generator Alexa Prize SocialBot that was deployed across four competition years. We also collect 35,623 candidate responses for the contexts. We annotate the responses with a four-level ABCD quality scheme that isolates conversational naturalness from topical relevance. We then extract twenty-five linguistic features that capture conversational properties of responses and contrast (A) responses that have a conversational style, from (B) responses that are topically relevant, but less conversational. We find that conversational responses are marked by other-directed engagement: question-asking, user engagement phrases, acknowledgment openings, and second-person reference, while responses that are merely relevant and topical are marked by self-oriented information delivery: opinion markers, formulaic openings, and hedges and emphasizers deployed in service of the system’s own assertions. We thus find that conversational style in this setting is best understood as a pragmatic orientation toward the user rather than toward the system’s own content, and that the system must be mixed-initiative to manifest a conversational style. We discuss what these findings imply for the design of Intelligent Virtual Agents.

Vrindavan Harrison, M. Walker · 0 citations
Preprint Aug 2026

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Hear2Act is introduced, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes that show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do not reliably carry it into action without an explicit intermediate representation.

Xinyi Liu, H. Nayyeri, Dilek Hakkani-Tur et al. · 1 citation
#natural language process... Preprint Aug 2026

You Know What I Mean: A Benchmark for Agentic Conversational Reference Grounding

The results show that CoRG remains challenging for current agents, even the best agent reaches only 67.0% success rate, leaving one third of references unresolved, and position CoRG as a concrete benchmark for studying how agents search, inspect, and verify information in realistic multi-tool environments.

Karen Fuchs, Uri Katz, Yoav Goldberg · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.