End-to-end losses based on speaker basis vectors and all-speaker hard   negative mining for speaker verification

Hee-Soo Heo; Jee-weon Jung; IL-Ho Yang; Sung-Hyun Yoon; Hye-jin Shim,; and Ha-Jin Yu

arXiv:1902.02455·eess.AS·July 18, 2019·1 cites

End-to-end losses based on speaker basis vectors and all-speaker hard negative mining for speaker verification

Hee-Soo Heo, Jee-weon Jung, IL-Ho Yang, Sung-Hyun Yoon, Hye-jin Shim,, and Ha-Jin Yu

PDF

Open Access

TL;DR

This paper introduces two novel end-to-end loss functions for speaker verification that leverage speaker basis vectors, enabling consideration of all speakers during training and improving inter-speaker variation and hard negative mining.

Contribution

The proposed loss functions utilize trainable speaker basis vectors to incorporate all speakers in training, enhancing speaker verification performance over traditional methods.

Findings

01

Improved speaker verification accuracy on VoxCeleb datasets.

02

Effective incorporation of all speakers in loss calculations.

03

Enhanced inter-speaker variation and hard negative mining capabilities.

Abstract

In recent years, speaker verification has primarily performed using deep neural networks that are trained to output embeddings from input features such as spectrograms or Mel-filterbank energies. Studies that design various loss functions, including metric learning have been widely explored. In this study, we propose two end-to-end loss functions for speaker verification using the concept of speaker bases, which are trainable parameters. One loss function is designed to further increase the inter-speaker variation, and the other is designed to conduct the identical concept with hard negative mining. Each speaker basis is designed to represent the corresponding speaker in the process of training deep neural networks. In contrast to the conventional loss functions that can consider only a limited number of speakers included in a mini-batch, the proposed loss functions can consider all the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Speech and Audio Processing · Music and Audio Processing

MethodsSoftmax