A Pre-training Framework that Encodes Noise Information for Speech   Quality Assessment

Subrina Sultana; Donald S. Williamson

arXiv:2411.04379·eess.AS·November 8, 2024·ICASSP

A Pre-training Framework that Encodes Noise Information for Speech Quality Assessment

Subrina Sultana, Donald S. Williamson

PDF

Open Access

TL;DR

This paper introduces a pre-training framework that encodes background noise information alongside speech features, enhancing speech quality assessment by leveraging noise cues often ignored by traditional SSL methods.

Contribution

The proposed framework uniquely combines supervised noise encoding with self-supervised speech embedding, improving perceptual speech quality estimation with fewer parameters.

Findings

01

Improved speech quality assessment performance.

02

Effective noise information encoding in representations.

03

Fewer parameters needed compared to baselines.

Abstract

Self-supervised learning (SSL) has grown in interest within the speech processing community, since it produces representations that are useful for many downstream tasks. SSL uses global and contextual methods to produce robust representations, where SSL even outperforms supervised models. Most self-supervised approaches, however, are limited to embedding information about, i.e., the phonemes, speaker identity, and emotion, into the extracted representations, where they become invariant to background sounds due to contrastive and auto-regressive learning. This is limiting because many downstream tasks leverage noise information to function accurately. Therefore, we propose a pre-training framework that learns information pertaining to background noise in a supervised manner, while jointly embedding speech information using a self-supervised strategy. We experiment with multiple encoders…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Vehicle Noise and Vibration Control · Music and Audio Processing