Comparative Analysis of Mel-Frequency Cepstral Coefficients and Wavelet   Based Audio Signal Processing for Emotion Detection and Mental Health   Assessment in Spoken Speech

Idoko Agbo; Dr Hoda El-Sayed; M.D Kamruzzan Sarker

arXiv:2412.10469·cs.SD·December 17, 2024

Comparative Analysis of Mel-Frequency Cepstral Coefficients and Wavelet Based Audio Signal Processing for Emotion Detection and Mental Health Assessment in Spoken Speech

Idoko Agbo, Dr Hoda El-Sayed, M.D Kamruzzan Sarker

PDF

Open Access

TL;DR

This study compares MFCC and wavelet features for emotion detection in speech using CNN and LSTM models, achieving up to 61% accuracy, and suggests future enhancements for mental health assessment.

Contribution

It introduces a comparative analysis of MFCC and wavelet features with CNN and LSTM models for emotion detection in speech, highlighting their effectiveness and areas for improvement.

Findings

01

CNN outperformed LSTM with 61% accuracy

02

Both models excelled in detecting surprise and anger

03

Distinct audio features like pitch and speed were influential

Abstract

The intersection of technology and mental health has spurred innovative approaches to assessing emotional well-being, particularly through computational techniques applied to audio data analysis. This study explores the application of Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) models on wavelet extracted features and Mel-frequency Cepstral Coefficients (MFCCs) for emotion detection from spoken speech. Data augmentation techniques, feature extraction, normalization, and model training were conducted to evaluate the models' performance in classifying emotional states. Results indicate that the CNN model achieved a higher accuracy of 61% compared to the LSTM model's accuracy of 56%. Both models demonstrated better performance in predicting specific emotions such as surprise and anger, leveraging distinct audio features like pitch and speed variations.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsEmotion and Mood Recognition

MethodsSPEED: Separable Pyramidal Pooling EncodEr-Decoder for Real-Time Monocular Depth Estimation on Low-Resource Settings · Sigmoid Activation · Tanh Activation · Long Short-Term Memory