Feature-informed Latent Space Regularization for Music Source Separation

Yun-Ning Hung; Alexander Lerch

arXiv:2203.09132·eess.AS·June 28, 2022

Feature-informed Latent Space Regularization for Music Source Separation

Yun-Ning Hung, Alexander Lerch

PDF

Open Access

TL;DR

This paper proposes transfer learning strategies that incorporate VGGish features into a music source separation model, improving performance without requiring additional annotations during training or inference.

Contribution

It introduces three novel approaches, including latent space regularization methods, to effectively integrate VGGish features into source separation models.

Findings

01

Improved evaluation metrics for music source separation.

02

Latent space regularization outperforms naive concatenation.

03

VGGish features enhance separation quality.

Abstract

The integration of additional side information to improve music source separation has been investigated numerous times, e.g., by adding features to the input or by adding learning targets in a multi-task learning scenario. These approaches, however, require additional annotations such as musical scores, instrument labels, etc. in training and possibly during inference. The available datasets for source separation do not usually provide these additional annotations. In this work, we explore transfer learning strategies to incorporate VGGish features with a state-of-the-art source separation model; VGGish features are known to be a very condensed representation of audio content and have been successfully used in many MIR tasks. We introduce three approaches to incorporate the features, including two latent space regularization methods and one naive concatenation method. Experimental…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Music and Audio Processing · Speech Recognition and Synthesis