Audio-Guided Fusion Techniques for Multimodal Emotion Analysis

Pujin Shi; Fei Gao

arXiv:2409.05007·cs.SD·September 10, 2024

Audio-Guided Fusion Techniques for Multimodal Emotion Analysis

Pujin Shi, Fei Gao

PDF

Open Access

TL;DR

This paper introduces an audio-guided fusion approach for multimodal emotion analysis, combining fine-tuned feature extractors, a novel transformer fusion mechanism, and self-supervised learning to improve sentiment classification accuracy.

Contribution

It presents a new Audio-Guided Transformer fusion mechanism and a semi-supervised learning strategy for multimodal emotion analysis, achieving competitive results in MER2024.

Findings

01

Achieved third place in MER-SEMI track

02

Effective fusion of audio, video, and text features

03

Improved sentiment classification performance

Abstract

In this paper, we propose a solution for the semi-supervised learning track (MER-SEMI) in MER2024. First, in order to enhance the performance of the feature extractor on sentiment classification tasks,we fine-tuned video and text feature extractors, specifically CLIP-vit-large and Baichuan-13B, using labeled data. This approach effectively preserves the original emotional information conveyed in the videos. Second, we propose an Audio-Guided Transformer (AGT) fusion mechanism, which leverages the robustness of Hubert-large, showing superior effectiveness in fusing both inter-channel and intra-channel information. Third, To enhance the accuracy of the model, we iteratively apply self-supervised learning by using high-confidence unlabeled data as pseudo-labels. Finally, through black-box probing, we discovered an imbalanced data distribution between the training and test sets. Therefore,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Emotion and Mood Recognition · Color perception and design

MethodsAttention Is All You Need · Byte Pair Encoding · Absolute Position Encodings · Softmax · Label Smoothing · Dropout · Layer Normalization · Position-Wise Feed-Forward Layer · Linear Layer · Adam