Improving Few-Shot Learning for Talking Face System with TTS Data   Augmentation

Qi Chen; Ziyang Ma; Tao Liu; Xu Tan; Qu Lu; Xie Chen; Kai Yu

arXiv:2303.05322·cs.SD·March 10, 2023·1 cites

Improving Few-Shot Learning for Talking Face System with TTS Data Augmentation

Qi Chen, Ziyang Ma, Tao Liu, Xu Tan, Qu Lu, Xie Chen, Kai Yu

PDF

Open Access 1 Repo

TL;DR

This paper introduces a TTS data augmentation method for improving few-shot learning in talking face systems, addressing data scarcity issues and enhancing synthesis quality through innovative alignment and feature extraction techniques.

Contribution

The paper proposes a novel TTS-based data augmentation approach combined with soft-DTW alignment and HuBERT features to enhance few-shot talking face synthesis.

Findings

01

Achieved 17% improvement in MSE score

02

Achieved 14% improvement in DTW score

03

Achieved 38% preference in user study

Abstract

Audio-driven talking face has attracted broad interest from academia and industry recently. However, data acquisition and labeling in audio-driven talking face are labor-intensive and costly. The lack of data resource results in poor synthesis effect. To alleviate this issue, we propose to use TTS (Text-To-Speech) for data augmentation to improve few-shot ability of the talking face system. The misalignment problem brought by the TTS audio is solved with the introduction of soft-DTW, which is first adopted in the talking face task. Moreover, features extracted by HuBERT are explored to utilize underlying information of audio, and found to be superior over other features. The proposed method achieves 17%, 14%, 38% dominance on MSE score, DTW score and user study preference repectively over the baseline model, which shows the effectiveness of improving few-shot learning for talking face…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

moon0316/t2a
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Face recognition and analysis