Multi-level Attention Fusion Network for Audio-visual Event Recognition

Mathilde Brousmiche; Jean Rouat; St\'ephane Dupont

arXiv:2106.06736·cs.CV·June 15, 2021

Multi-level Attention Fusion Network for Audio-visual Event Recognition

Mathilde Brousmiche, Jean Rouat, St\'ephane Dupont

PDF

Open Access 1 Repo

TL;DR

The paper introduces MAFnet, a neural network architecture that dynamically fuses audio and visual data at multiple levels for improved event recognition in videos, inspired by neuroscience principles.

Contribution

It proposes a novel multi-level attention fusion network that adaptively combines audio-visual modalities at different processing stages for better classification accuracy.

Findings

01

Improves accuracy on AVE, UCF51, and Kinetics-Sounds datasets.

02

Effectively highlights relevant modality and time windows.

03

Demonstrates the benefit of multi-level attention in multimodal fusion.

Abstract

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion network (MAFnet), an architecture that can dynamically fuse visual and audio information for event recognition. Inspired by prior studies in neuroscience, we couple both modalities at different levels of visual and audio paths. Furthermore, the network dynamically highlights a modality at a given time window relevant to classify events. Experimental results in AVE (Audio-Visual Event), UCF51, and Kinetics-Sounds datasets show that the approach can effectively improve the accuracy in audio-visual event classification. Code is available at: https://github.com/numediart/MAFnet

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

numediart/MAFnet
tfOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Speech and Audio Processing · Digital Media Forensic Detection