Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!

Prateek Verma

arXiv:2309.08751·cs.SD·May 8, 2025

Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!

Prateek Verma

PDF

Open Access

TL;DR

This paper demonstrates that combining diverse handcrafted audio feature embeddings with end-to-end learned representations significantly improves sound classification performance over using end-to-end models alone.

Contribution

It introduces a method to integrate domain-specific handcrafted audio embeddings with end-to-end models, achieving superior classification results.

Findings

01

Handcrafted embeddings alone do not outperform end-to-end models.

02

Combining handcrafted and end-to-end embeddings improves accuracy.

03

The approach surpasses traditional end-to-end training in audio classification.

Abstract

With the advent of modern AI architectures, a shift has happened towards end-to-end architectures. This pivot has led to neural architectures being trained without domain-specific biases/knowledge, optimized according to the task. We in this paper, learn audio embeddings via diverse feature representations, in this case, domain-specific. For the case of audio classification over hundreds of categories of sound, we learn robust separate embeddings for diverse audio properties such as pitch, timbre, and neural representation, along with also learning it via an end-to-end architecture. We observe handcrafted embeddings, e.g., pitch and timbre-based, although on their own, are not able to beat a fully end-to-end representation, yet adding these together with end-to-end embedding helps us, significantly improve performance. This work would pave the way to bring some domain expertise with…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Speech and Audio Processing · Music Technology and Sound Studies