Autovocoder: Fast Waveform Generation from a Learned Speech   Representation using Differentiable Digital Signal Processing

Jacob J Webber; Cassia Valentini-Botinhao; Evelyn Williams; Gustav Eje; Henter; Simon King

arXiv:2211.06989·cs.SD·May 25, 2023·1 cites

Autovocoder: Fast Waveform Generation from a Learned Speech Representation using Differentiable Digital Signal Processing

Jacob J Webber, Cassia Valentini-Botinhao, Evelyn Williams, Gustav Eje, Henter, Simon King

PDF

Open Access 1 Repo

TL;DR

The paper introduces an autovocoder that uses learned representations and differentiable digital signal processing to generate speech waveforms efficiently, achieving faster synthesis with comparable quality to existing neural vocoders.

Contribution

It proposes a novel autovocoder framework that replaces traditional mel-spectrograms with learned representations and employs differentiable DSP for rapid waveform synthesis.

Findings

01

Generates waveforms 5 times faster than Griffin-Lim

02

Achieves 14 times faster synthesis than HiFi-GAN

03

Perceptual tests show comparable speech quality to HiFi-GAN

Abstract

Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation. A mel-spectrogram is extracted from the waveform by a simple, fast DSP operation, but generating a high-quality waveform from a mel-spectrogram requires computationally expensive machine learning: a neural vocoder. Our proposed ``autovocoder'' reverses this arrangement. We use machine learning to obtain a representation that replaces the mel-spectrogram, and that can be inverted back to a waveform using simple, fast operations including a differentiable implementation of the inverse STFT. The autovocoder generates a waveform 5 times faster than the DSP-based Griffin-Lim algorithm, and 14 times faster than the neural vocoder HiFi-GAN. We provide perceptual listening test results to confirm that the speech is of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

hcy71o/autovocoder
pytorch

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Speech Recognition and Synthesis · Advanced Data Compression Techniques