Prosodic Phrase Alignment for Machine Dubbing

Alp \"Oktem; Mireia Farr\'us; Antonio Bonafonte

arXiv:1908.07226·cs.CL·August 21, 2019

Prosodic Phrase Alignment for Machine Dubbing

Alp \"Oktem, Mireia Farr\'us, Antonio Bonafonte

PDF

1 Repo

TL;DR

This paper presents a neural attention-based method for prosodic phrase synchronization in machine dubbing, improving lip-sync accuracy and speech rate matching compared to traditional approaches.

Contribution

It introduces a novel approach leveraging neural attention mechanisms to enhance prosodic alignment in machine dubbing, addressing a key challenge in audiovisual translation.

Findings

01

Achieved speech rate ratios comparable to professional dubbing

02

Improved lip-sync quality for long dialogue lines

03

Demonstrated effectiveness of attention-based phrasing in dubbing

Abstract

Dubbing is a type of audiovisual translation where dialogues are translated and enacted so that they give the impression that the media is in the target language. It requires a careful alignment of dubbed recordings with the lip movements of performers in order to achieve visual coherence. In this paper, we deal with the specific problem of prosodic phrase synchronization within the framework of machine dubbing. Our methodology exploits the attention mechanism output in neural machine translation to find plausible phrasing for the translated dialogue lines and then uses them to condition their synthesis. Our initial work in this field records comparable speech rate ratio to professional dubbing translation, and improvement in terms of lip-syncing of long dialogue lines.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

alpoktem/MachineDub
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.