An Investigation of End-to-End Models for Robust Speech Recognition

Archiki Prasad; Preethi Jyothi; Rajbabu Velmurugan

arXiv:2102.06237·eess.AS·February 15, 2021

An Investigation of End-to-End Models for Robust Speech Recognition

Archiki Prasad, Preethi Jyothi, Rajbabu Velmurugan

PDF

1 Repo

TL;DR

This paper systematically compares speech enhancement and model adaptation techniques for robust end-to-end speech recognition, revealing that the effectiveness depends on noise type and highlighting the trade-offs involved.

Contribution

It provides a comprehensive comparison of enhancement-based and model-based adaptation methods for end-to-end robust ASR, a gap in prior research.

Findings

01

Adversarial learning performs best on certain noise types but degrades clean speech WER.

02

A new speech enhancement technique outperforms model-based methods on stationary noise.

03

Knowledge of noise type influences the choice of adaptation technique.

Abstract

End-to-end models for robust automatic speech recognition (ASR) have not been sufficiently well-explored in prior work. With end-to-end models, one could choose to preprocess the input speech using speech enhancement techniques and train the model using enhanced speech. Another alternative is to pass the noisy speech as input and modify the model architecture to adapt to noisy speech. A systematic comparison of these two approaches for end-to-end robust ASR has not been attempted before. We address this gap and present a detailed comparison of speech enhancement-based techniques and three different model-based adaptation techniques covering data augmentation, multi-task learning, and adversarial learning for robust ASR. While adversarial learning is the best-performing technique on certain noise types, it comes at the cost of degrading clean speech WER. On other relatively stationary…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

archiki/Robust-E2E-ASR
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.