MMSD-Net: Towards Multi-modal Stuttering Detection

Liangyu Nie; Sudarsana Reddy Kadiri; and Ruchit Agrawal

arXiv:2407.11492·cs.SD·July 17, 2024

MMSD-Net: Towards Multi-modal Stuttering Detection

Liangyu Nie, Sudarsana Reddy Kadiri, and Ruchit Agrawal

PDF

Open Access

TL;DR

This paper introduces MMSD-Net, a multi-modal neural framework that combines speech and visual signals to improve automatic stuttering detection, achieving significant performance gains over uni-modal methods.

Contribution

MMSD-Net is the first multi-modal neural approach for stuttering detection, integrating visual signals to enhance detection accuracy.

Findings

01

Incorporating visual signals improves detection performance.

02

MMSD-Net outperforms uni-modal approaches by 2-17% in F1-score.

03

Multi-modal approach demonstrates significant potential for speech disorder detection.

Abstract

Stuttering is a common speech impediment that is caused by irregular disruptions in speech production, affecting over 70 million people across the world. Standard automatic speech processing tools do not take speech ailments into account and are thereby not able to generate meaningful results when presented with stuttered speech as input. The automatic detection of stuttering is an integral step towards building efficient, context-aware speech processing systems. While previous approaches explore both statistical and neural approaches for stuttering detection, all of these methods are uni-modal in nature. This paper presents MMSD-Net, the first multi-modal neural framework for stuttering detection. Experiments and results demonstrate that incorporating the visual signal significantly aids stuttering detection, and our model yields an improvement of 2-17% in the F1-score over existing…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStuttering Research and Treatment · Text Readability and Simplification · Employee Welfare and Language Studies