Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation

Akam Rahimi; Triantafyllos Afouras; Andrew Zisserman

arXiv:2501.01518·eess.AS·January 6, 2025

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation

Akam Rahimi, Triantafyllos Afouras, Andrew Zisserman

PDF

Open Access

TL;DR

This paper introduces a Transformer-based multi-modal speech separation framework that leverages visual and textual cues, demonstrating robustness to synchronization issues and achieving state-of-the-art results on benchmark datasets.

Contribution

It presents a novel unified multi-modal speech separation model using Transformers, incorporating textual and visual cues, and handling asynchronous inputs effectively.

Findings

01

State-of-the-art performance on LRS2 and LRS3 datasets

02

Robustness to audio-visual synchronization offsets

03

Effective fusion of visual and textual modalities

Abstract

The goal of this paper is speech separation and enhancement in multi-speaker and noisy environments using a combination of different modalities. Previous works have shown good performance when conditioning on temporal or static visual evidence such as synchronised lip movements or face identity. In this paper, we present a unified framework for multi-modal speech separation and enhancement based on synchronous or asynchronous cues. To that end we make the following contributions: (i) we design a modern Transformer-based architecture tailored to fuse different modalities to solve the speech separation task in the raw waveform domain; (ii) we propose conditioning on the textual content of a sentence alone or in combination with visual information; (iii) we demonstrate the robustness of our model to audio-visual synchronisation offsets; and, (iv) we obtain state-of-the-art performance on…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Phonetics and Phonology Research · Speech Recognition and Synthesis