Population Based Training for Data Augmentation and Regularization in   Speech Recognition

Daniel Haziza; J\'er\'emy Rapin; Gabriel Synnaeve

arXiv:2010.03899·cs.CL·October 9, 2020

Population Based Training for Data Augmentation and Regularization in Speech Recognition

Daniel Haziza, J\'er\'emy Rapin, Gabriel Synnaeve

PDF

Open Access

TL;DR

This paper introduces population-based training to optimize data augmentation and regularization schedules in speech recognition, leading to significant performance improvements and reduced experimental effort.

Contribution

It demonstrates the effectiveness of population-based training for dynamic hyperparameter optimization in speech recognition, simplifying the process and improving accuracy.

Findings

01

8% relative WER improvement over baseline

02

Achieved 5.18% WER on LibriSpeech test-other

03

Effective optimization of SpecAugment and dropout schedules

Abstract

Varying data augmentation policies and regularization over the course of optimization has led to performance improvements over using fixed values. We show that population based training is a useful tool to continuously search those hyperparameters, within a fixed budget. This greatly simplifies the experimental burden and computational cost of finding such optimal schedules. We experiment in speech recognition by optimizing SpecAugment this way, as well as dropout. It compares favorably to a baseline that does not change those hyperparameters over the course of training, with an 8% relative WER improvement. We obtain 5.18% word error rate on LibriSpeech's test-other.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Music and Audio Processing · Topic Modeling

MethodsPopulation Based Training