A Closer Look at Neural Codec Resynthesis: Bridging the Gap between   Codec and Waveform Generation

Alexander H. Liu; Qirui Wang; Yuan Gong; James Glass

arXiv:2410.22448·eess.AS·October 31, 2024

A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation

Alexander H. Liu, Qirui Wang, Yuan Gong, James Glass

PDF

Open Access

TL;DR

This paper investigates methods to improve waveform resynthesis from neural audio codec tokens, highlighting the impact of learning targets and introducing a Schr"odinger Bridge approach for enhanced speech quality.

Contribution

It introduces a novel Schr"odinger Bridge-based resynthesis method and analyzes how different strategies influence audio quality and perception.

Findings

01

Schr"odinger Bridge improves waveform reconstruction quality.

02

Choice of learning target significantly affects audio perception.

03

Different resynthesis strategies impact both machine and human evaluations.

Abstract

Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence of tokens, i.e., discrete embeddings. The discrete and low-frequency nature of neural codecs introduced a new way to generate speech with token-based models. As these tokens encode information at various levels of granularity, from coarse to fine, most existing works focus on how to better generate the coarse tokens. In this paper, we focus on an equally important but often overlooked question: How can we better resynthesize the waveform from coarse tokens? We point out that both the choice of learning target and resynthesis approach have a dramatic impact on the generated audio quality. Specifically, we study two different strategies based on token prediction and regression, and introduce a new method based…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNeural Networks and Applications

MethodsSoftmax · Attention Is All You Need · Focus