Text-aware and Context-aware Expressive Audiobook Speech Synthesis

Dake Guo; Xinfa Zhu; Liumeng Xue; Yongmao Zhang; Wenjie Tian; Lei Xie

arXiv:2406.05672·eess.AS·June 13, 2024·Interspeech

Text-aware and Context-aware Expressive Audiobook Speech Synthesis

Dake Guo, Xinfa Zhu, Liumeng Xue, Yongmao Zhang, Wenjie Tian, Lei Xie

PDF

Open Access

TL;DR

This paper introduces a novel text-aware and context-aware style modeling approach for expressive audiobook speech synthesis, capturing diverse narration styles without manual labels, thereby enhancing naturalness and expressiveness.

Contribution

It proposes a new style modeling framework that combines contrastive learning and cross-sentence context encoding, integrated into existing TTS models for improved audiobook narration.

Findings

01

Effectively captures diverse narration styles.

02

Improves naturalness and expressiveness of synthesized speech.

03

Enhances coherence across sentences in audiobook synthesis.

Abstract

Recent advances in text-to-speech have significantly improved the expressiveness of synthetic speech. However, a major challenge remains in generating speech that captures the diverse styles exhibited by professional narrators in audiobooks without relying on manually labeled data or reference speech. To address this problem, we propose a text-aware and context-aware(TACA) style modeling approach for expressive audiobook speech synthesis. We first establish a text-aware style space to cover diverse styles via contrastive learning with the supervision of the speech style. Meanwhile, we adopt a context encoder to incorporate cross-sentence information and the style embedding obtained from text. Finally, we introduce the context encoder to two typical TTS models, VITS-based TTS and language model-based TTS. Experimental results demonstrate that our proposed approach can effectively capture…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Speech and Audio Processing · Music and Audio Processing

MethodsContrastive Learning