DeepFry: Identifying Vocal Fry Using Deep Neural Networks

Bronya R. Chernyak; Talia Ben Simon; Yael Segal; Jeremy Steffman,; Eleanor Chodroff; Jennifer S. Cole; Joseph Keshet

arXiv:2203.17019·eess.AS·June 28, 2022

DeepFry: Identifying Vocal Fry Using Deep Neural Networks

Bronya R. Chernyak, Talia Ben Simon, Yael Segal, Jeremy Steffman,, Eleanor Chodroff, Jennifer S. Cole, Joseph Keshet

PDF

1 Repo

TL;DR

This paper introduces a deep learning model that accurately detects vocal fry in speech, addressing challenges posed by irregular glottal vibrations and improving recognition systems for languages with prevalent creaky voice.

Contribution

It presents a novel encoder-classifier deep neural network that learns from raw waveforms and refines creak detection using auxiliary voice features, outperforming previous methods.

Findings

01

Improved recall and F1 scores on unseen data

02

Effective use of raw waveform and auxiliary features

03

Enhanced detection of creaky voice in American English

Abstract

Vocal fry or creaky voice refers to a voice quality characterized by irregular glottal opening and low pitch. It occurs in diverse languages and is prevalent in American English, where it is used not only to mark phrase finality, but also sociolinguistic factors and affect. Due to its irregular periodicity, creaky voice challenges automatic speech processing and recognition systems, particularly for languages where creak is frequently used. This paper proposes a deep learning model to detect creaky voice in fluent speech. The model is composed of an encoder and a classifier trained together. The encoder takes the raw waveform and learns a representation using a convolutional neural network. The classifier is implemented as a multi-headed fully-connected network trained to detect creaky voice, voicing, and pitch, where the last two are used to refine creak prediction. The model is…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

bronichern/deepfry
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Voice and Speech Disorders · Phonetics and Phonology Research

Methods7 Fastest Ways to Call American Airlines Reservations Number (USA Guide)