Speaker Change Detection for Transformer Transducer ASR

Jian Wu; Zhuo Chen; Min Hu; Xiong Xiao; Jinyu Li

arXiv:2302.08549·eess.AS·February 20, 2023

Speaker Change Detection for Transformer Transducer ASR

Jian Wu, Zhuo Chen, Min Hu, Xiong Xiao, Jinyu Li

PDF

Open Access

TL;DR

This paper introduces a new framework for speaker change detection that enhances performance by adding an SCD module on top of Transformer Transducer ASR, enabling independent optimization and improving F1 scores.

Contribution

The paper proposes a novel SCD framework built on Transformer Transducer ASR allowing independent optimization and improved speaker change detection accuracy.

Findings

01

Significant F1 score improvements on LibriCSS and Microsoft datasets.

02

SCD module operates independently without degrading ASR performance.

03

Two variants of the SCD network effectively estimate speaker change probabilities.

Abstract

Speaker change detection (SCD) is an important feature that improves the readability of the recognized words from an automatic speech recognition (ASR) system by breaking the word sequence into paragraphs at speaker change points. Existing SCD solutions either require additional ensemble for the time based decisions and recognized word sequences, or implement a tight integration between ASR and SCD, limiting the potential optimum performance for both tasks. To address these issues, we propose a novel framework for the SCD task, where an additional SCD module is built on top of an existing Transformer Transducer ASR (TT-ASR) network. Two variants of the SCD network are explored in this framework that naturally estimate speaker change probability for each word, while allowing the ASR and SCD to have independent optimization scheme for the best performance. Experiments show that our…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Speech and Audio Processing · Music and Audio Processing

MethodsMulti-Head Attention · Attention Is All You Need · Linear Layer · Label Smoothing · Dense Connections · Absolute Position Encodings · Adam · Position-Wise Feed-Forward Layer · Dropout · Byte Pair Encoding