SpeakerStew: Scaling to Many Languages with a Triaged Multilingual   Text-Dependent and Text-Independent Speaker Verification System

Roza Chojnacka; Jason Pelecanos; Quan Wang; Ignacio Lopez Moreno

arXiv:2104.02125·eess.AS·June 17, 2021

SpeakerStew: Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification System

Roza Chojnacka, Jason Pelecanos, Quan Wang, Ignacio Lopez Moreno

PDF

1 Repo 1 Models

TL;DR

SpeakerStew introduces a scalable multilingual speaker verification system that combines data pooling and a triage mechanism to reduce computational costs and latency across 46 languages.

Contribution

The paper presents the first large-scale speaker verification system covering 46 languages, utilizing a novel triage approach to optimize performance and efficiency.

Findings

01

Training on multiple languages improves generalization to unseen languages.

02

The triage mechanism reduces computational calls by 73% and latency by 59%.

03

Performance remains robust with no worse EER than the baseline.

Abstract

In this paper, we describe SpeakerStew - a hybrid system to perform speaker verification on 46 languages. Two core ideas were explored in this system: (1) Pooling training data of different languages together for multilingual generalization and reducing development cycles; (2) A novel triage mechanism between text-dependent and text-independent models to reduce runtime cost and expected latency. To the best of our knowledge, this is the first study of speaker verification systems at the scale of 46 languages. The problem is framed from the perspective of using a smart speaker device with interactions consisting of a wake-up keyword (text-dependent) followed by a speech query (text-independent). Experimental evidence suggests that training on multiple languages can generalize to unseen varieties while maintaining performance on seen varieties. We also found that it can reduce…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

google/speaker-id/tree/master/lingvo
noneOfficial

Models

🤗
tflite-hub/conformer-speaker-encoder
model· 77 dl· ♡ 5
77 dl♡ 5

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.