VoxCeleb Enrichment for Age and Gender Recognition

Khaled Hechmi; Trung Ngo Trong; Ville Hautamaki; Tomi Kinnunen

arXiv:2109.13510·cs.LG·December 21, 2021

VoxCeleb Enrichment for Age and Gender Recognition

Khaled Hechmi, Trung Ngo Trong, Ville Hautamaki, Tomi Kinnunen

PDF

1 Repo

TL;DR

This paper enriches the VoxCeleb dataset with age and gender labels, and evaluates various models for recognizing these attributes from speech, highlighting challenges in age estimation.

Contribution

It provides new age and gender annotations for VoxCeleb and systematically compares multiple features and classifiers for attribute recognition.

Findings

01

Best gender recognition F1-score of 0.9829 with logistic regression.

02

Lowest age MAE of 9.443 years with ridge regression.

03

Identifies potential mislabels in original VoxCeleb data.

Abstract

VoxCeleb datasets are widely used in speaker recognition studies. Our work serves two purposes. First, we provide speaker age labels and (an alternative) annotation of speaker gender. Second, we demonstrate the use of this metadata by constructing age and gender recognition models with different features and classifiers. We query different celebrity databases and apply consensus rules to derive age and gender labels. We also compare the original VoxCeleb gender labels with our labels to identify records that might be mislabeled in the original VoxCeleb data. On modeling side, we design a comprehensive study of multiple features and models for recognizing gender and age. Our best system, using i-vector features, achieved an F1-score of 0.9829 for gender recognition task using logistic regression, and the lowest mean absolute error (MAE) in age regression, 9.443 years, is obtained with…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

hechmik/voxceleb_enrichment_age_gender
tfOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.