MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols

Yuhao Du; Qianwei Huang; Guo Zhu; Zhanchen Dai; Shunian Chen; Qiming Zhu; Le Pan; Minghao Chen; Yuhao Zhang; Li Zhou; Benyou Wang; and Haizhou Li

arXiv:2508.18240·cs.CL·September 16, 2025

MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols

Yuhao Du, Qianwei Huang, Guo Zhu, Zhanchen Dai, Shunian Chen, Qiming Zhu, Le Pan, Minghao Chen, Yuhao Zhang, Li Zhou, Benyou Wang, and Haizhou Li

PDF

1 Datasets

TL;DR

MTalk-Bench introduces a comprehensive multi-turn speech-to-speech evaluation framework that assesses models across semantic, paralinguistic, and ambient sound dimensions using dual evaluation methods, revealing current strengths and limitations.

Contribution

This work presents MTalk-Bench, a novel benchmark with dual evaluation protocols for multi-turn S2S models across multiple dimensions, addressing gaps in existing assessment frameworks.

Findings

01

Models excel at semantic processing but struggle with paralinguistic and ambient sounds.

02

Increasing response length improves coherence but reduces efficiency.

03

Modality-aware, task-specific models outperform brute-force scaling.

Abstract

The rapid advancement of speech-to-speech (S2S) large language models (LLMs) has significantly improved real-time spoken interaction. However, current evaluation frameworks remain inadequate for assessing performance in complex, multi-turn dialogues. To address this, we introduce MTalk-Bench, a multi-turn S2S benchmark covering three core dimensions: Semantic Information, Paralinguistic Information, and Ambient Sound. Each dimension includes nine realistic scenarios, along with targeted tasks to assess specific capabilities such as reasoning. Our dual-method evaluation framework combines Arena-style evaluation (pairwise comparison) and Rubrics-based evaluation (absolute scoring) for relative and absolute assessment. The benchmark includes both model and human outputs, evaluated by human evaluators and LLMs. Experimental results reveal two sets of findings. Overall performance of S2S…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Datasets

FreedomIntelligence/MTalk-Bench
dataset· 31 dl
31 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.