Independent Deeply Learned Tensor Analysis for Determined Audio Source   Separation

Naoki Narisawa; Rintaro Ikeshita; Norihiro Takamune; Daichi Kitamura,; Tomohiko Nakamura; Hiroshi Saruwatari; Tomohiro Nakatani

arXiv:2106.05529·cs.SD·June 11, 2021

Independent Deeply Learned Tensor Analysis for Determined Audio Source Separation

Naoki Narisawa, Rintaro Ikeshita, Norihiro Takamune, Daichi Kitamura,, Tomohiko Nakamura, Hiroshi Saruwatari, Tomohiro Nakatani

PDF

Open Access

TL;DR

This paper introduces a supervised deep neural network approach for audio source separation that models frequency covariance matrices to better capture nonstationary signals, outperforming previous methods.

Contribution

It proposes a novel FCM model combining diagonal and rank-1 matrices, and uses two DNNs for power spectrum and signal estimation, improving separation performance.

Findings

01

Higher separation performance than IDLMA

02

Effective modeling of nonstationary signals

03

Flexible FCM model capturing dynamics

Abstract

We address the determined audio source separation problem in the time-frequency domain. In independent deeply learned matrix analysis (IDLMA), it is assumed that the inter-frequency correlation of each source spectrum is zero, which is inappropriate for modeling nonstationary signals such as music signals. To account for the correlation between frequencies, independent positive semidefinite tensor analysis has been proposed. This unsupervised (blind) method, however, severely restrict the structure of frequency covariance matrices (FCMs) to reduce the number of model parameters. As an extension of these conventional approaches, we here propose a supervised method that models FCMs using deep neural networks (DNNs). It is difficult to directly infer FCMs using DNNs. Therefore, we also propose a new FCM model represented as a convex combination of a diagonal FCM and a rank-1 FCM. Our FCM…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsBlind Source Separation Techniques · Speech and Audio Processing · Advanced Adaptive Filtering Techniques