Multi-View Attention Transfer for Efficient Speech Enhancement

Wooseok Shin; Hyun Joon Park; Jin Sob Kim; Byung Hoon Lee; Sung Won; Han

arXiv:2208.10367·cs.SD·November 1, 2022·1 cites

Multi-View Attention Transfer for Efficient Speech Enhancement

Wooseok Shin, Hyun Joon Park, Jin Sob Kim, Byung Hoon Lee, Sung Won, Han

PDF

Open Access

TL;DR

This paper introduces multi-view attention transfer (MV-AT), a feature-based knowledge distillation method for efficient, low-complexity speech enhancement models that maintain high performance in the time domain.

Contribution

It proposes a novel MV-AT method that transfers multi-view knowledge without extra parameters, improving lightweight models for speech enhancement.

Findings

01

Significant reduction in parameters and FLOPs with maintained performance.

02

Consistent performance improvements across various model sizes.

03

Effective knowledge transfer demonstrated on Valentini and DNS datasets.

Abstract

Recent deep learning models have achieved high performance in speech enhancement; however, it is still challenging to obtain a fast and low-complexity model without significant performance degradation. Previous knowledge distillation studies on speech enhancement could not solve this problem because their output distillation methods do not fit the speech enhancement task in some aspects. In this study, we propose multi-view attention transfer (MV-AT), a feature-based distillation, to obtain efficient speech enhancement models in the time domain. Based on the multi-view features extraction model, MV-AT transfers multi-view knowledge of the teacher network to the student network without additional parameters. The experimental results show that the proposed method consistently improved the performance of student models of various sizes on the Valentini and deep noise suppression (DNS)…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Indoor and Outdoor Localization Technologies · Hand Gesture Recognition Systems

MethodsKnowledge Distillation