Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive   Conversational Speech Synthesis

Rui Liu; Zhenqi Jia; Feilong Bao; Haizhou Li

arXiv:2501.06467·cs.CL·January 14, 2025

Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis

Rui Liu, Zhenqi Jia, Feilong Bao, Haizhou Li

PDF

1 Repo

TL;DR

This paper introduces RADKA-CSS, a retrieval-augmented method that leverages stored dialogue knowledge to improve the expressiveness and style alignment of conversational speech synthesis.

Contribution

The paper proposes a novel retrieval-augmented dialogue knowledge aggregation scheme for expressive CSS, incorporating multi-attribute retrieval and multi-source style knowledge integration.

Findings

01

RADKA-CSS outperforms baseline models in expressiveness.

02

Effective retrieval of similar dialogues enhances speech style consistency.

03

The approach demonstrates significant improvements in both objective and subjective evaluations.

Abstract

Conversational speech synthesis (CSS) aims to take the current dialogue (CD) history as a reference to synthesize expressive speech that aligns with the conversational style. Unlike CD, stored dialogue (SD) contains preserved dialogue fragments from earlier stages of user-agent interaction, which include style expression knowledge relevant to scenarios similar to those in CD. Note that this knowledge plays a significant role in enabling the agent to synthesize expressive conversational speech that generates empathetic feedback. However, prior research has overlooked this aspect. To address this issue, we propose a novel Retrieval-Augmented Dialogue Knowledge Aggregation scheme for expressive CSS, termed RADKA-CSS, which includes three main components: 1) To effectively retrieve dialogues from SD that are similar to CD in terms of both semantic and style. First, we build a stored…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

coder-jzq/radka-css
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.