Approach to Learning Generalized Audio Representation Through Batch   Embedding Covariance Regularization and Constant-Q Transforms

Ankit Shah; Shuyi Chen; Kejun Zhou; Yue Chen; Bhiksha Raj

arXiv:2303.03591·cs.SD·March 8, 2023·1 cites

Approach to Learning Generalized Audio Representation Through Batch Embedding Covariance Regularization and Constant-Q Transforms

Ankit Shah, Shuyi Chen, Kejun Zhou, Yue Chen, Bhiksha Raj

PDF

Open Access

TL;DR

This paper introduces a novel regularization technique and evaluates different audio preprocessing methods to enhance general-purpose audio embeddings, demonstrating improved dispersion and performance across diverse tasks.

Contribution

It proposes Batch Embedding Covariance Regularization (BECR) and compares Constant-Q Transform with STFT, advancing audio representation learning.

Findings

01

BECR leads to more dispersed embeddings.

02

BECR improves the PaSST model without extra complexity.

03

STFT preprocessing outperforms CQT across tasks.

Abstract

General-purpose embedding is highly desirable for few-shot even zero-shot learning in many application scenarios, including audio tasks. In order to understand representations better, we conducted a thorough error analysis and visualization of HEAR 2021 submission results. Inspired by the analysis, this work experiments with different front-end audio preprocessing methods, including Constant-Q Transform (CQT) and Short-time Fourier transform (STFT), and proposes a Batch Embedding Covariance Regularization (BECR) term to uncover a more holistic simulation of the frequency information received by the human auditory system. We tested the models on the suite of HEAR 2021 tasks, which encompass a broad category of tasks. Preliminary results show (1) the proposed BECR can incur a more dispersed embedding on the test set, (2) BECR improves the PaSST model without extra computation complexity,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Speech and Audio Processing · Speech Recognition and Synthesis

MethodsTest