Discrete Audio Representation as an Alternative to Mel-Spectrograms for   Speaker and Speech Recognition

Krishna C. Puvvada; Nithin Rao Koluguri; Kunal Dhawan; Jagadeesh; Balam; Boris Ginsburg

arXiv:2309.10922·eess.AS·September 21, 2023·1 cites

Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition

Krishna C. Puvvada, Nithin Rao Koluguri, Kunal Dhawan, Jagadeesh, Balam, Boris Ginsburg

PDF

Open Access

TL;DR

This paper evaluates compression-based discrete audio tokens as an alternative to mel-spectrograms for speaker verification, diarization, and speech recognition, showing competitive performance, robustness, and high compression ratios.

Contribution

It provides a comprehensive comparison of compression-based audio tokens with mel-spectrograms across multiple speech tasks, highlighting their potential and limitations.

Findings

01

Models trained on audio tokens are within 1% of mel-spectrogram performance.

02

Audio tokens are robust to out-of-domain narrowband data.

03

Achieves 20x compression with minimal performance loss.

Abstract

Discrete audio representation, aka audio tokenization, has seen renewed interest driven by its potential to facilitate the application of text language modeling approaches in audio domain. To this end, various compression and representation-learning based tokenization schemes have been proposed. However, there is limited investigation into the performance of compression-based audio tokens compared to well-established mel-spectrogram features across various speaker and speech related tasks. In this paper, we evaluate compression based audio tokens on three tasks: Speaker Verification, Diarization and (Multi-lingual) Speech Recognition. Our findings indicate that (i) the models trained on audio tokens perform competitively, on average within $1%$ of mel-spectrogram features for all the tasks considered, and do not surpass them yet. (ii) these models exhibit robustness for out-of-domain…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Speech Recognition and Synthesis · Music and Audio Processing