Utterance Weighted Multi-Dilation Temporal Convolutional Networks for   Monaural Speech Dereverberation

William Ravenscroft; Stefan Goetze; Thomas Hain

arXiv:2205.08455·cs.SD·July 26, 2022

Utterance Weighted Multi-Dilation Temporal Convolutional Networks for Monaural Speech Dereverberation

William Ravenscroft, Stefan Goetze, Thomas Hain

PDF

Open Access 1 Repo

TL;DR

This paper introduces a weighted multi-dilation depthwise-separable convolution for TCNs, enhancing speech dereverberation by dynamically focusing on local and global information, leading to improved performance and parameter efficiency.

Contribution

It proposes a novel weighted multi-dilation convolution to improve TCNs for speech dereverberation, outperforming standard TCNs with fewer parameters.

Findings

01

WD-TCN outperforms baseline TCN in SISDR by up to 0.55 dB

02

The proposed method is more parameter efficient than increasing model size

03

Achieves 12.26 dB SISDR on WHAMR dataset

Abstract

Speech dereverberation is an important stage in many speech technology applications. Recent work in this area has been dominated by deep neural network models. Temporal convolutional networks (TCNs) are deep learning models that have been proposed for sequence modelling in the task of dereverberating speech. In this work a weighted multi-dilation depthwise-separable convolution is proposed to replace standard depthwise-separable convolutions in TCN models. This proposed convolution enables the TCN to dynamically focus on more or less local information in its receptive field at each convolutional block in the network. It is shown that this weighted multi-dilation temporal convolutional network (WD-TCN) consistently outperforms the TCN across various model configurations and using the WD-TCN model is a more parameter efficient method to improve the performance of the model than increasing…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

jwr1995/wd-tcn
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Speech Recognition and Synthesis · Phonetics and Phonology Research

MethodsConvolution