Enhanced Neural Beamformer with Spatial Information for Target Speech   Extraction

Aoqi Guo; Junnan Wu; Peng Gao; Wenbo Zhu; Qinwen Guo; Dazhi Gao and; Yujun Wang

arXiv:2306.15942·cs.SD·June 29, 2023

Enhanced Neural Beamformer with Spatial Information for Target Speech Extraction

Aoqi Guo, Junnan Wu, Peng Gao, Wenbo Zhu, Qinwen Guo, Dazhi Gao and, Yujun Wang

PDF

Open Access

TL;DR

This paper introduces an enhanced neural beamformer that leverages spatial information and advanced neural network structures to improve target speech extraction accuracy in noisy environments.

Contribution

It proposes a novel target speech extraction network combining UNet-TCN and multi-head cross-attention to better utilize spatial cues for speech separation.

Findings

01

Significant improvement in speech separation accuracy.

02

Effective utilization of spatial information via cross-attention.

03

Enhanced neural beamformer performance demonstrated through experiments.

Abstract

Recently, deep learning-based beamforming algorithms have shown promising performance in target speech extraction tasks. However, most systems do not fully utilize spatial information. In this paper, we propose a target speech extraction network that utilizes spatial information to enhance the performance of neural beamformer. To achieve this, we first use the UNet-TCN structure to model input features and improve the estimation accuracy of the speech pre-separation module by avoiding information loss caused by direct dimensionality reduction in other models. Furthermore, we introduce a multi-head cross-attention mechanism that enhances the neural beamformer's perception of spatial information by making full use of the spatial information received by the array. Experimental results demonstrate that our approach, which incorporates a more reasonable target mask estimation network and a…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Speech Recognition and Synthesis · Music and Audio Processing