Acoustic and linguistic representations for speech continuous emotion   recognition in call center conversations

Manon Macary; Marie Tahon; Yannick Est\`eve; Daniel Luzzati

arXiv:2310.04481·eess.AS·October 10, 2023

Acoustic and linguistic representations for speech continuous emotion recognition in call center conversations

Manon Macary, Marie Tahon, Yannick Est\`eve, Daniel Luzzati

PDF

Open Access

TL;DR

This study investigates continuous emotion recognition in call center conversations, emphasizing the dominance of linguistic features over acoustic ones, and explores transfer learning with pre-trained speech representations to improve performance.

Contribution

It demonstrates the effectiveness of pre-trained linguistic models like CamemBERT for emotion prediction and analyzes the robustness of fusion approaches and variability factors in annotations.

Findings

01

Linguistic content is the main contributor to emotion prediction.

02

Pre-trained linguistic features significantly outperform acoustic features.

03

Fusion of modalities offers limited additional benefit.

Abstract

The goal of our research is to automatically retrieve the satisfaction and the frustration in real-life call-center conversations. This study focuses an industrial application in which the customer satisfaction is continuously tracked down to improve customer services. To compensate the lack of large annotated emotional databases, we explore the use of pre-trained speech representations as a form of transfer learning towards AlloSat corpus. Moreover, several studies have pointed out that emotion can be detected not only in speech but also in facial trait, in biological response or in textual information. In the context of telephone conversations, we can break down the audio information into acoustic and linguistic by using the speech signal and its transcription. Our experiments confirms the large gain in performance obtained with the use of pre-trained features. Surprisingly, we found…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsEmotion and Mood Recognition · Speech and Audio Processing