Enhancing into the codec: Noise Robust Speech Coding with   Vector-Quantized Autoencoders

Jonah Casebeer; Vinjai Vale; Umut Isik; Jean-Marc Valin; Ritwik Giri,; Arvindh Krishnaswamy

arXiv:2102.06610·eess.AS·February 15, 2021

Enhancing into the codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders

Jonah Casebeer, Vinjai Vale, Umut Isik, Jean-Marc Valin, Ritwik Giri,, Arvindh Krishnaswamy

PDF

TL;DR

This paper introduces noise-robust speech coding using vector-quantized autoencoders, improving performance in noisy environments and outperforming models trained solely on clean speech.

Contribution

It develops compressor-enhancer autoencoders based on VQ-VAE with WaveRNN decoders, enhancing noise robustness in speech coding.

Findings

01

Enhanced noise robustness in speech coding models

02

Compressor-enhancer models outperform pure compressor models in noisy conditions

03

Models trained on both clean and noisy speech perform better

Abstract

Audio codecs based on discretized neural autoencoders have recently been developed and shown to provide significantly higher compression levels for comparable quality speech output. However, these models are tightly coupled with speech content, and produce unintended outputs in noisy conditions. Based on VQ-VAE autoencoders with WaveRNN decoders, we develop compressor-enhancer encoders and accompanying decoders, and show that they operate well in noisy conditions. We also observe that a compressor-enhancer model performs better on clean speech inputs than a compressor model trained only on clean speech.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

Methods*Communicated@Fast*How Do I Communicate to Expedia? · Sigmoid Activation · Tanh Activation · Softmax · WaveRNN · VQ-VAE