SIG-VC: A Speaker Information Guided Zero-shot Voice Conversion System   for Both Human Beings and Machines

Haozhe Zhang; Zexin Cai; Xiaoyi Qin; Ming Li

arXiv:2111.03811·cs.SD·April 4, 2023

SIG-VC: A Speaker Information Guided Zero-shot Voice Conversion System for Both Human Beings and Machines

Haozhe Zhang, Zexin Cai, Xiaoyi Qin, Ming Li

PDF

Open Access 1 Repo

TL;DR

This paper introduces SIG-VC, a zero-shot voice conversion system that effectively disentangles speaker and content information, improving conversion quality and robustness under extreme conditions for both humans and machines.

Contribution

The paper presents a novel framework for zero-shot voice conversion that enhances speaker-content disentanglement and maintains voice cloning performance with added speaker control.

Findings

01

Significantly reduces the trade-off in zero-shot voice conversion

02

Achieves high spoofing power against speaker verification systems

03

Demonstrates effectiveness through subjective and objective evaluations

Abstract

Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people's attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice conversion. We aim to obtain intermediate representations for speaker-content disentanglement of speech to better remove speaker information and get pure content information. Accordingly, our proposed framework contains a module that removes the speaker information from the acoustic feature of the source speaker. Moreover, speaker information control is added to our system to maintain the voice cloning performance. The proposed system is evaluated by subjective and objective metrics. Results show that our proposed system significantly reduces the trade-off problem in zero-shot voice conversion, while it also manages to have high spoofing power to the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

HaydenCaffrey/SIG-VC
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Speech and Audio Processing · Music and Audio Processing