Advancing RNN Transducer Technology for Speech Recognition

George Saon; Zoltan Tueske; Daniel Bolanos; Brian Kingsbury

arXiv:2103.09935·cs.CL·March 19, 2021

Advancing RNN Transducer Technology for Speech Recognition

George Saon, Zoltan Tueske, Daniel Bolanos, Brian Kingsbury

PDF

TL;DR

This paper presents new techniques for RNN Transducers, including architectural innovations and adaptation methods, significantly reducing word error rates across multiple speech recognition tasks.

Contribution

It introduces a novel multiplicative integration in the joint network and explores speaker adaptation, language model fusion, and training strategies for improved RNN-T performance.

Findings

01

Achieved 5.9% WER on Switchboard test set.

02

Reduced WER by 12.5% on CallHome.

03

Attained 12.7% WER on Italian test set.

Abstract

We investigate a set of techniques for RNN Transducers (RNN-Ts) that were instrumental in lowering the word error rate on three different tasks (Switchboard 300 hours, conversational Spanish 780 hours and conversational Italian 900 hours). The techniques pertain to architectural changes, speaker adaptation, language model fusion, model combination and general training recipe. First, we introduce a novel multiplicative integration of the encoder and prediction network vectors in the joint network (as opposed to additive). Second, we discuss the applicability of i-vector speaker adaptation to RNN-Ts in conjunction with data perturbation. Third, we explore the effectiveness of the recently proposed density ratio language model fusion for these tasks. Last but not least, we describe the other components of our training recipe and their effect on recognition performance. We report a 5.9% and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.