Acoustic-to-articulatory inversion for dysarthric speech: Are   pre-trained self-supervised representations favorable?

Sarthak Kumar Maharana; Krishna Kamal Adidam; Shoumik Nandi; Ajitesh; Srivastava

arXiv:2309.01108·eess.AS·February 13, 2024

Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?

Sarthak Kumar Maharana, Krishna Kamal Adidam, Shoumik Nandi, Ajitesh, Srivastava

PDF

Open Access

TL;DR

This study evaluates the effectiveness of pre-trained self-supervised learning representations for acoustic-to-articulatory inversion in dysarthric speech, showing that SSL features outperform traditional MFCCs, especially when fine-tuned.

Contribution

It demonstrates that SSL representations like DeCoAR, wav2vec, and APC improve AAI performance on dysarthric speech, particularly with fine-tuning, in low-resource scenarios.

Findings

01

DeCoAR with fine-tuning improves correlation by ~4.56% for patients.

02

SSL features outperform MFCCs in both seen and unseen cases.

03

SSL models trained on reconstruction or prediction tasks perform well for dysarthric speech.

Abstract

Acoustic-to-articulatory inversion (AAI) involves mapping from the acoustic to the articulatory space. Signal-processing features like the MFCCs, have been widely used for the AAI task. For subjects with dysarthric speech, AAI is challenging because of an imprecise and indistinct pronunciation. In this work, we perform AAI for dysarthric speech using representations from pre-trained self-supervised learning (SSL) models. We demonstrate the impact of different pre-trained features on this challenging AAI task, at low-resource conditions. In addition, we also condition x-vectors to the extracted SSL features to train a BLSTM network. In the seen case, we experiment with three AAI training schemes (subject-specific, pooled, and fine-tuned). The results, consistent across training schemes, reveal that DeCoAR, in the fine-tuned scheme, achieves a relative improvement of the Pearson…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsVoice and Speech Disorders · Phonetics and Phonology Research · Speech Recognition and Synthesis