Robust speaker recognition using unsupervised adversarial invariance

Raghuveer Peri; Monisankha Pal; Arindam Jati; Krishna Somandepalli,; Shrikanth Narayanan

arXiv:1911.00940·eess.AS·November 5, 2019

Robust speaker recognition using unsupervised adversarial invariance

Raghuveer Peri, Monisankha Pal, Arindam Jati, Krishna Somandepalli,, Shrikanth Narayanan

PDF

1 Repo

TL;DR

This paper introduces an unsupervised adversarial invariance approach to extract robust speaker embeddings, significantly improving speaker recognition and diarization performance in challenging acoustic environments.

Contribution

It presents a novel unsupervised adversarial invariance architecture that disentangles speaker information from acoustic variability without supervision.

Findings

01

Outperforms baseline in challenging acoustic scenarios

02

Achieves 36% relative improvement in diarization error rate

03

Enhances robustness of speaker embeddings for verification and clustering

Abstract

In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervised adversarial invariance architecture to train a network that maps speaker embeddings extracted using a pre-trained model onto two lower dimensional embedding spaces. The embedding spaces are learnt to disentangle speaker-discriminative information from all other information present in the audio recordings, without supervision about the acoustic conditions. We analyze the robustness of the proposed embeddings to various sources of variability present in the signal for speaker verification and unsupervised clustering tasks on a large-scale speaker recognition corpus. Our analyses show that the proposed system substantially outperforms the baseline in a variety of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

rperi/speaker-embeddings-UAI-inference
tf

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.