Multi-encoder attention-based architectures for sound recognition with   partial visual assistance

Wim Boes; Hugo Van hamme

arXiv:2209.12826·eess.AS·October 11, 2022

Multi-encoder attention-based architectures for sound recognition with partial visual assistance

Wim Boes, Hugo Van hamme

PDF

Open Access

TL;DR

This paper introduces a multi-encoder attention-based framework that integrates partial visual information into sound recognition models, enhancing performance even when visual data is intermittently unavailable.

Contribution

The study presents a novel multi-encoder architecture that incorporates partial visual features into deep learning sound recognition systems, addressing data availability issues.

Findings

01

Improved accuracy in audio tagging and sound event detection

02

Effective handling of missing visual data during inference

03

Insights into limitations of multi-encoder visual integration

Abstract

Large-scale sound recognition data sets typically consist of acoustic recordings obtained from multimedia libraries. As a consequence, modalities other than audio can often be exploited to improve the outputs of models designed for associated tasks. Frequently, however, not all contents are available for all samples of such a collection: For example, the original material may have been removed from the source platform at some point, and therefore, non-auditory features can no longer be acquired. We demonstrate that a multi-encoder framework can be employed to deal with this issue by applying this method to attention-based deep learning systems, which are currently part of the state of the art in the domain of sound recognition. More specifically, we show that the proposed model extension can successfully be utilized to incorporate partially available visual information into the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Speech and Audio Processing · Noise Effects and Management