Joint Speaker Counting, Speech Recognition, and Speaker Identification   for Overlapped Speech of Any Number of Speakers

Naoyuki Kanda; Yashesh Gaur; Xiaofei Wang; Zhong Meng; Zhuo Chen,; Tianyan Zhou; Takuya Yoshioka

arXiv:2006.10930·eess.AS·August 11, 2020·5 cites

Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers

Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen,, Tianyan Zhou, Takuya Yoshioka

PDF

Open Access

TL;DR

This paper introduces an end-to-end model that simultaneously performs speaker counting, speech recognition, and speaker identification on overlapped speech, improving accuracy over separate methods.

Contribution

The authors extend serialized output training with a speaker inventory and joint optimization, enabling unified processing of overlapped speech with multiple speakers.

Findings

01

Significantly improved speaker-attributed word error rate

02

Effective joint modeling of speech recognition and speaker ID

03

Outperforms baseline methods on LibriSpeech

Abstract

We propose an end-to-end speaker-attributed automatic speech recognition model that unifies speaker counting, speech recognition, and speaker identification on monaural overlapped speech. Our model is built on serialized output training (SOT) with attention-based encoder-decoder, a recently proposed method for recognizing overlapped speech comprising an arbitrary number of speakers. We extend SOT by introducing a speaker inventory as an auxiliary input to produce speaker labels as well as multi-speaker transcriptions. All model parameters are optimized by speaker-attributed maximum mutual information criterion, which represents a joint probability for overlapped speech recognition and speaker identification. Experiments on LibriSpeech corpus show that our proposed method achieves significantly better speaker-attributed word error rate than the baseline that separately performs…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Speech and Audio Processing · Music and Audio Processing