Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 9230-9241· 0 citations· 10 references
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in multimodal understanding and reasoning by integrating linguistic and visual information, while benchmarks facilitate iterative model improvement by evaluating their performance and analyzing their limitations. However, existing dialogue-based multimodal benchmarks do not fully reflect the characteristics of real-world interactions, as they often construct a single, lengthy user utterance to provide all requirements or treat visual information as static even in multi-turn conversations. To address these limitations, we propose the Multi-turn Multimodal Interactive Dialogue (MMID) Benchmark, where user requirements are incrementally conveyed across turns and images are interleaved with text throughout the conversation to enable dynamic multimodal interaction. With this design, MMID enables comprehensive evaluation of the Perception, Memorization, and Reasoning abilities of MLLMs. Furthermore, while most tasks adopt a multiple-choice question format, each incorrect option is mapped to fine-grained error types, enabling an analysis of model strengths and weaknesses beyond coarse-grained performance comparison. MMID reveals MLLMs perform well with text-based input but degrade with images, requiring improved leverage fine-grained visual cues. Our benchmarks and detailed descriptions are available at https://github.com/KUNLP/MMID.
This work analyzes real chatbot failures to identify six recurring mechanisms and defines six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding, which shows that Hy-MultiTurn is broadly challenging.
Eileen Ye, Ji-Hua Tao, Yao-Ming Li et al.· arXiv.org· 0 citations
This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents, and concludes with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures.
Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti et al.· 2 citations
This work proposes M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs and evaluates models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior.
Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding. We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded generation; (2) DAILYDIALOG-FA, 6.6K dialogues annotated for dialogue acts and emotions; and (3) PLAYDIAL-FA, 2.1K theatrical dialogues with sentiment labels. While LLMs assist data construction, every dialogue undergoes multi-stage review and revision by native Farsi speakers, and only the final human-approved dialogues are released. Experiments with six LLAMA and MISTRAL models show that LoRA substantially improves dialogue generation while requiring only 25-50% of the training data to recover over 90% of the final performance gains. Across classification tasks, FABERT achieves the best dialogue-act performance, LORA-MISTRAL-7B performs best on emotion recognition, and MISTRAL-24B achieves the highest sentiment score. Human evaluation and independent external validation demonstrate the reliability of the benchmark, while comparisons with GPT-4.1 as an LLM judge reveal that automatic metrics substantially overestimate dialogue quality. Zero-shot evaluation with frontier LLMs further shows that TalkFa remains a challenging benchmark. We will release all datasets, annotation guidelines, code, and checkpoints.
Neda Jamshidi, Kamyar Zeinalipour, F. Akbari et al.· 0 citations
The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance as general voice assistants, their performance in specialized domains remains underexplored, particularly in scientific areas. Scientific interactions introduce formidable challenges, involving rare technical terminology, spoken norms of abbreviations, and the natural verbalization of symbolic special expressions. In this paper, we introduce S$^3$-Bench, a systematic evaluation framework covering 10 major disciplines, consisting of a Knowledge set for speech question-answering and a Dialogue set for multi-turn progressive interactions with simulated user agents. By decomposing a complete atomic turn into stages of speech recognition, perception, knowledge utilization with reasoning, and response pronunciation, we systematically characterize the common challenges and performance tradeoffs of existing approaches. Furthermore, experiments on multi-turn interactions reveal persistent limitations in user adaptation and the generation of accurate, comprehensive, and efficient responses.
He-Yang Liu, Jia-Yi Huang, Wen Xiao et al.· 0 citations
Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much referential information gestures and their combination with speech carry in multimodal dialogue under different partner visibility conditions. % We build models that identify the intended referent in a video-mediated referential communication game based on either the speech transcript, the skeletal representation of gesture, or both modalities. Our results show that gesture alone is predictive of the intended referent and that multimodal fusion is most beneficial when the transcript-based model is uncertain. Training-only alignment of learned representations with the referent image further improves the fusion model performance. % In a comparison with human interaction data, we further see pragmatic effects of interlocutor visibility on gesture production and informativeness as well as an entrainment effect in speech and multimodal, but not gesture, performance across rounds of repeated interaction. We thus make contributions to the technical modelling of multimodal information in human dialogue and the analysis of human interaction data via trained model representations.
E. Ghaleb, Hugh Mee Wong, Kristina Kobrock· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.