Skip to content

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

Jul 2026 · arXiv.org · Vol abs/2607.29125 · 0 citations · 31 references
Computer Science

TL;DR

This work proposes M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs and evaluates models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior.

Abstract

Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations remain a challenge. In addition, existing benchmarks provide limited coverage of languages and dialogue domains. We propose M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs. M3-DuplexBench supports English and Japanese and covers both casual conversation and multi-turn question answering. In addition, we evaluate models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior. Experiments with recent FDSDSs reveal model-specific turn-taking characteristics, clear performance gaps across languages and domains, and mixed effects of dialogue context.

View source

Similar papers

Preprint Sep 2026

From Metrics to Natural Dialogue: French Full-Duplex Benchmark for Spoken Dialogue Models

Full-duplex spoken dialogue models aim to make voice agents more natural by allowing them to listen, speak, pause, and respond during ongoing conversation. However, it is not clear whether full-duplex benchmarks behave the same way when models are evaluated in a different language. To investigate this, we introduce a French full-duplex benchmark (FDB) with two variants, CALLFC-FDB for Canadian French and MEDIA-FDB for European French, and compare them with an English FDB. Built from real spoken resources, these benchmarks evaluate key full-duplex skills, including pause handling, turn-taking, backchannels, and interruptions. Beyond introducing French FDBs, we evaluate human--human conversations with FDB metrics to better understand the values these metrics take in real-world dialogue. Our analysis reveals that most timing-based metrics behave similarly across languages, while content-based evaluation degrades under language mismatch. We also find a trade-off between optimizing benchmark metrics and preserving conversational naturalness.

Hamid Soltani, Gilles Boulianne · 0 citations
#natural language process... Preprint Aug 2026

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection, finds end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in backchannel-dense interaction styles.

Freeman Jiang, Ramon Sanabria, Soham Deshmukh et al. · 2 citations
#artificial intelligence Preprint Aug 2026

Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework

This work proposes LPS-TC, a Lightweight Proactive Speech Turn Controller for plug-and-play integration, and introduces a two-tier evaluation scheme that assesses both chunk-level timing precision and turn-level interaction quality under realistic streaming constraints.

Tianrui Pan, Qinglin Zhang, Chong Deng et al. · 0 citations
Review Aug 2026

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents, and concludes with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures.

Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti et al. · 2 citations
Jul 2026

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

This work analyzes real chatbot failures to identify six recurring mechanisms and defines six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding, which shows that Hy-MultiTurn is broadly challenging.

Eileen Ye, Ji-Hua Tao, Yao-Ming Li et al. · 0 citations
Book Open access Aug 2026

MMID: Multi-turn Multimodal Interactive Dialogue Benchmark

The proposed Multi-turn Multimodal Interactive Dialogue (MMID) Benchmark enables comprehensive evaluation of the Perception, Memorization, and Reasoning abilities of MLLMs, and reveals MLLMs perform well with text-based input but degrade with images, requiring improved leverage fine-grained visual cues.

Seulgi Kim, Juoh Sun, Sumin Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.