Block-Online Guided Source Separation

Shota Horiguchi; Yusuke Fujita; Kenji Nagamatsu

arXiv:2011.07791·eess.AS·November 17, 2020·1 cites

Block-Online Guided Source Separation

Shota Horiguchi, Yusuke Fujita, Kenji Nagamatsu

PDF

Open Access

TL;DR

This paper introduces a block-online guided source separation algorithm that reduces computation time and latency, enabling real-time multi-talker speech separation using diarization information.

Contribution

The proposed algorithm enables real-time speech separation by processing blocks with context, reducing computation and latency compared to offline methods.

Findings

01

Achieved nearly the same separation performance as offline GSS

02

32x faster computation suitable for real-time applications

03

Effective in multi-talker scenarios with diarization info.

Abstract

We propose a block-online algorithm of guided source separation (GSS). GSS is a speech separation method that uses diarization information to update parameters of the generative model of observation signals. Previous studies have shown that GSS performs well in multi-talker scenarios. However, it requires a large amount of calculation time, which is an obstacle to the deployment of online applications. It is also a problem that the offline GSS is an utterance-wise algorithm so that it produces latency according to the length of the utterance. With the proposed algorithm, block-wise input samples and corresponding time annotations are concatenated with those in the preceding context and used to update the parameters. Using the context enables the algorithm to estimate time-frequency masks accurately only from one iteration of optimization for each block, and its latency does not depend…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Speech Recognition and Synthesis · Advanced Adaptive Filtering Techniques