This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents, and concludes with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures.
Abstract
Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
The proposed Multi-turn Multimodal Interactive Dialogue (MMID) Benchmark enables comprehensive evaluation of the Perception, Memorization, and Reasoning abilities of MLLMs, and reveals MLLMs perform well with text-based input but degrade with images, requiring improved leverage fine-grained visual cues.
Seulgi Kim, Juoh Sun, Sumin Kim et al.· Proceedings of the 32nd ACM...· 0 citations
This work proposes M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs and evaluates models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior.
Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interactions due to the exponentially greater conversational complexity. For voice agents to integrate seamlessly into human group dynamics, they must not only generate contextually appropriate responses but also demonstrate a nuanced understanding of open turn-taking. To address this gap, we introduce Multiparty Bench (MP-Bench), the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts. MP-Bench assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness. Additionally, we incorporate comprehension-based question-answering tasks as a complementary evaluation. By benchmarking 12 voice agents, we find that real-time voice agents stay at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking, exposing an open challenge for real-time voice agents under multiparty scenario.
Yi-Jen Shih, S. Kuan, Guan-Ting Lin et al.· 0 citations
Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks. Their utility depends on continuity: agents must retrieve and use earlier evidence across dialogue, files, and workspace state. However, agents can generate plausible answers even when access to that history has degraded, causing outcome-only evaluation to overestimate true evidence use. We present MIRAGE (Multimodal Interaction Retrieval, Attribution, and Grounding Evaluation), a controlled study of historical evidence use under conversation-state variation in multimodal personal agents. MIRAGE holds evidence objects, questions, and scoring fixed while varying only conversation state, and evaluates whether an agent can determine answerability, recover the correct source, and answer from it. Across seven frontier and open-weight multimodal backbones, we find that: 1) pre-compaction depth and post-compaction continuation form distinct, non-monotonic failure regimes rather than a single degradation curve; 2) open-weight models rely heavily on context continuity and are reluctant to spontaneously switch to tool-mediated retrieval when provenance fails; and 3) retrieval pressure improves source attribution in deep pre-compaction states for tool-compliant models, but consistently regresses after compaction, where stored evidence has already degraded. These findings show that historical evidence use should be evaluated under state variation, rather than inferred from outcome-only correctness.
Yu Liu, Wen-Xiao Zhang, Cheng Hu et al.· 0 citations
TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection, finds end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in backchannel-dense interaction styles.
Freeman Jiang, Ramon Sanabria, Soham Deshmukh et al.· 2 citations
Intelligent Virtual Agents need to be able to participate in extended dialogue interactions while maintaining a conversational style. We discover the elements of conversational style in open-domain dialogues by analyzing the features that distinguish conversational system responses from responses that are merely topically relevant. We first collect 7,751 dialogue contexts from live human conversations with a multi-generator Alexa Prize SocialBot that was deployed across four competition years. We also collect 35,623 candidate responses for the contexts. We annotate the responses with a four-level ABCD quality scheme that isolates conversational naturalness from topical relevance. We then extract twenty-five linguistic features that capture conversational properties of responses and contrast (A) responses that have a conversational style, from (B) responses that are topically relevant, but less conversational. We find that conversational responses are marked by other-directed engagement: question-asking, user engagement phrases, acknowledgment openings, and second-person reference, while responses that are merely relevant and topical are marked by self-oriented information delivery: opinion markers, formulaic openings, and hedges and emphasizers deployed in service of the system’s own assertions. We thus find that conversational style in this setting is best understood as a pragmatic orientation toward the user rather than toward the system’s own content, and that the system must be mixed-initiative to manifest a conversational style. We discuss what these findings imply for the design of Intelligent Virtual Agents.
Vrindavan Harrison, M. Walker· Proceedings of the 26th ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.