Voice Conversion by Cascading Automatic Speech Recognition and   Text-to-Speech Synthesis with Prosody Transfer

Jing-Xuan Zhang; Li-Juan Liu; Yan-Nian Chen; Ya-Jun Hu; Yuan Jiang,; Zhen-Hua Ling; Li-Rong Dai

arXiv:2009.01475·eess.AS·September 4, 2020·5 cites

Voice Conversion by Cascading Automatic Speech Recognition and Text-to-Speech Synthesis with Prosody Transfer

Jing-Xuan Zhang, Li-Juan Liu, Yan-Nian Chen, Ya-Jun Hu, Yuan Jiang,, Zhen-Hua Ling, Li-Rong Dai

PDF

Open Access

TL;DR

This paper introduces a voice conversion method combining ASR and TTS with prosody transfer, achieving high naturalness and speaker similarity by transferring prosody features during synthesis.

Contribution

It proposes a novel prosody transfer approach using a prosody encoder in a cascading ASR-TTS system for improved voice conversion quality.

Findings

01

Achieved top naturalness and similarity in Voice Conversion Challenge 2020

02

Demonstrated effective prosody transfer improves voice conversion

03

Validated the method's effectiveness through experiments

Abstract

With the development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) technique, it's intuitive to construct a voice conversion system by cascading an ASR and TTS system. In this paper, we present a ASR-TTS method for voice conversion, which used iFLYTEK ASR engine to transcribe the source speech into text and a Transformer TTS model with WaveNet vocoder to synthesize the converted speech from the decoded text. For the TTS model, we proposed to use a prosody code to describe the prosody information other than text and speaker information contained in speech. A prosody encoder is used to extract the prosody code. During conversion, the source prosody is transferred to converted speech by conditioning the Transformer TTS model with its code. Experiments were conducted to demonstrate the effectiveness of our proposed method. Our system also obtained the best…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis

MethodsMulti-Head Attention · Attention Is All You Need · Linear Layer · Absolute Position Encodings · Position-Wise Feed-Forward Layer · Softmax · Dropout · Adam · Layer Normalization · Label Smoothing