Cross-modal Speaker Verification and Recognition: A Multilingual   Perspective

Muhammad Saad Saeed; Shah Nawaz; Pietro Morerio; Arif Mahmood; Ignazio; Gallo; Muhammad Haroon Yousaf; and Alessio Del Bue

arXiv:2004.13780·cs.CV·April 23, 2021·1 cites

Cross-modal Speaker Verification and Recognition: A Multilingual Perspective

Muhammad Saad Saeed, Shah Nawaz, Pietro Morerio, Arif Mahmood, Ignazio, Gallo, Muhammad Haroon Yousaf, and Alessio Del Bue

PDF

Open Access

TL;DR

This paper investigates whether face-voice association and speaker recognition are language-independent by introducing a multilingual dataset and conducting experiments to evaluate cross-lingual biometric matching.

Contribution

It introduces a new multilingual audio-visual dataset and explores the challenges of cross-lingual face-voice association and speaker recognition, addressing key questions in multilingual biometric systems.

Findings

01

Face-voice association shows partial language independence.

02

Speaker recognition performance varies across languages.

03

Multilingual challenges significantly impact biometric system effectiveness.

Abstract

Recent years have seen a surge in finding association between faces and voices within a cross-modal biometric application along with speaker recognition. Inspired from this, we introduce a challenging task in establishing association between faces and voices across multiple languages spoken by the same set of persons. The aim of this paper is to answer two closely related questions: "Is face-voice association language independent?" and "Can a speaker be recognised irrespective of the spoken language?". These two questions are very important to understand effectiveness and to boost development of multilingual biometric systems. To answer them, we collected a Multilingual Audio-Visual dataset, containing human speech clips of $154$ identities with $3$ language annotations extracted from various videos uploaded online. Extensive experiments on the three splits of the proposed dataset have…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Face recognition and analysis · Music and Audio Processing