Contrastive Predictive Coding Supported Factorized Variational   Autoencoder for Unsupervised Learning of Disentangled Speech Representations

Janek Ebbers; Michael Kuhlmann; Tobias Cord-Landwehr; Reinhold; Haeb-Umbach

arXiv:2005.12963·eess.AS·March 12, 2021

Contrastive Predictive Coding Supported Factorized Variational Autoencoder for Unsupervised Learning of Disentangled Speech Representations

Janek Ebbers, Michael Kuhlmann, Tobias Cord-Landwehr, Reinhold, Haeb-Umbach

PDF

TL;DR

This paper introduces a novel unsupervised method using contrastive predictive coding within a variational autoencoder to disentangle style and content in speech, improving robustness and performance without requiring labeled data.

Contribution

It proposes a fully convolutional variational autoencoder with adversarial contrastive predictive coding for unsupervised speech disentanglement, outperforming existing methods.

Findings

01

Effective separation of speaker and content traits.

02

Enhanced robustness of content representations against train-test mismatch.

03

Competitive performance in unsupervised speaker-content disentanglement.

Abstract

In this work we address disentanglement of style and content in speech signals. We propose a fully convolutional variational autoencoder employing two encoders: a content encoder and a style encoder. To foster disentanglement, we propose adversarial contrastive predictive coding. This new disentanglement method does neither need parallel data nor any supervision. We show that the proposed technique is capable of separating speaker and content traits into the two different representations and show competitive speaker-content disentanglement performance compared to other unsupervised approaches. We further demonstrate an increased robustness of the content representation against a train-test mismatch compared to spectral features, when used for phone recognition.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.