SpeechSplit 2.0: Unsupervised speech disentanglement for voice   conversion Without tuning autoencoder Bottlenecks

Chak Ho Chan; Kaizhi Qian; Yang Zhang; Mark Hasegawa-Johnson

arXiv:2203.14156·eess.AS·March 29, 2022

SpeechSplit 2.0: Unsupervised speech disentanglement for voice conversion Without tuning autoencoder Bottlenecks

Chak Ho Chan, Kaizhi Qian, Yang Zhang, Mark Hasegawa-Johnson

PDF

Open Access 1 Repo

TL;DR

SpeechSplit 2.0 introduces an unsupervised speech disentanglement method that avoids autoencoder bottleneck tuning by using signal processing techniques, resulting in more robust voice conversion performance.

Contribution

It replaces autoencoder bottleneck tuning with signal processing constraints, enhancing robustness and simplifying the disentanglement process.

Findings

01

Achieves comparable speech disentanglement performance to SpeechSplit

02

Demonstrates superior robustness to bottleneck size variations

03

Maintains effective aspect-specific voice conversion

Abstract

SpeechSplit can perform aspect-specific voice conversion by disentangling speech into content, rhythm, pitch, and timbre using multiple autoencoders in an unsupervised manner. However, SpeechSplit requires careful tuning of the autoencoder bottlenecks, which can be time-consuming and less robust. This paper proposes SpeechSplit 2.0, which constrains the information flow of the speech component to be disentangled on the autoencoder input using efficient signal processing methods instead of bottleneck tuning. Evaluation results show that SpeechSplit 2.0 achieves comparable performance to SpeechSplit in speech disentanglement and superior robustness to the bottleneck size variations. Our code is available at https://github.com/biggytruck/SpeechSplit2.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

biggytruck/speechsplit2
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Speech and Audio Processing