Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

Srikanth Korse; Mohamed Elminshawi; Emanuel A. P. Habets; Srikanth Raj Chetupalli

arXiv:2507.06566·eess.AS·July 10, 2025·ICASSP

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

Srikanth Korse, Mohamed Elminshawi, Emanuel A. P. Habets, Srikanth Raj Chetupalli

PDF

Open Access

TL;DR

This paper introduces a modality dropout training strategy for multi-modal target speaker extraction, enhancing robustness against modality dominance and improving performance when one modality is unavailable.

Contribution

The study proposes modality dropout training as a novel approach to improve robustness in multi-modal speaker extraction systems, outperforming standard and multi-task training methods.

Findings

01

MDT outperforms standard and MTT strategies across experiments.

02

Models with MDT are less affected by normalization layer choices.

03

MDT-trained systems are robust to using speech as enrollment signals.

Abstract

The primary goal of multi-modal TSE (MTSE) is to extract a target speaker from a speech mixture using complementary information from different modalities, such as audio enrolment and visual feeds corresponding to the target speaker. MTSE systems are expected to perform well even when one of the modalities is unavailable. In practice, the systems often suffer from modality dominance, where one of the modalities outweighs the others, thereby limiting robustness. Our study investigates training strategies and the effect of architectural choices, particularly the normalization layers, in yielding a robust MTSE system in both non-causal and causal configurations. In particular, we propose the use of modality dropout training (MDT) as a superior strategy to standard and multi-task training (MTT) strategies. Experiments conducted on two-speaker mixtures from the LRS3 dataset show the MDT…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Speech and Audio Processing · Phonetics and Phonology Research