Multi-input Architecture and Disentangled Representation Learning for   Multi-dimensional Modeling of Music Similarity

Sebastian Ribecky; Jakob Abe{\ss}er; Hanna Lukashevich

arXiv:2111.01710·eess.AS·November 3, 2021

Multi-input Architecture and Disentangled Representation Learning for Multi-dimensional Modeling of Music Similarity

Sebastian Ribecky, Jakob Abe{\ss}er, Hanna Lukashevich

PDF

Open Access

TL;DR

This paper introduces a multi-input neural network that processes various audio features to model and analyze multi-dimensional music similarity, outperforming existing methods and providing insights into how different musical factors influence similarity perception.

Contribution

The paper presents a novel multi-input architecture that directly models human multi-dimensional music similarity using disentangled features from multiple audio representations.

Findings

01

Outperforms state-of-the-art similarity prediction methods

02

Effectively models multiple musical dimensions such as genre, mood, and tempo

03

Provides a multi-dimensional analysis of factors influencing music similarity

Abstract

In the context of music information retrieval, similarity-based approaches are useful for a variety of tasks that benefit from a query-by-example scenario. Music however, naturally decomposes into a set of semantically meaningful factors of variation. Current representation learning strategies pursue the disentanglement of such factors from deep representations, resulting in highly interpretable models. This allows the modeling of music similarity perception, which is highly subjective and multi-dimensional. While the focus of prior work is on metadata driven notions of similarity, we suggest to directly model the human notion of multi-dimensional music similarity. To achieve this, we propose a multi-input deep neural network architecture, which simultaneously processes mel-spectrogram, CENS-chromagram and tempogram in order to extract informative features for the different disentangled…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Neuroscience and Music Perception · Music Technology and Sound Studies