Spectrogram-channels u-net: a source separation model viewing each   channel as the spectrogram of each source

Jaehoon Oh; Duyeon Kim; and Se-Young Yun

arXiv:1810.11520·cs.SD·October 31, 2018·1 cites

Spectrogram-channels u-net: a source separation model viewing each channel as the spectrogram of each source

Jaehoon Oh, Duyeon Kim, and Se-Young Yun

PDF

Open Access 1 Repo

TL;DR

This paper introduces Spectrogram-Channels U-Net, a novel source separation model that treats each output channel as a spectrogram of a separated source, achieving state-of-the-art results in singing voice and multi-instrument separation.

Contribution

It presents a new spectrogram-based U-Net model with a volume-balancing loss function, adaptable to various source separation tasks.

Findings

01

Achieved state-of-the-art separation performance

02

Effective for both singing voice and multi-instrument separation

03

Introduced a volume-balancing loss function

Abstract

Sound source separation has attracted attention from Music Information Retrieval(MIR) researchers, since it is related to many MIR tasks such as automatic lyric transcription, singer identification, and voice conversion. In this paper, we propose an intuitive spectrogram-based model for source separation by adapting U-Net. We call it Spectrogram-Channels U-Net, which means each channel of the output corresponds to the spectrogram of separated source itself. The proposed model can be used for not only singing voice separation but also multi-instrument separation by changing only the number of output channels. In addition, we propose a loss function that balances volumes between different sources. Finally, we yield performance that is state-of-the-art on both separation tasks.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

lucas-dunker/stem-separator-amt
pytorch

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Music and Audio Processing · Speech Recognition and Synthesis

MethodsConcatenated Skip Connection · *Communicated@Fast*How Do I Communicate to Expedia? · Max Pooling · Convolution · U-Net