# Multimodal music information processing and retrieval: survey and future   challenges

**Authors:** Federico Simonetta, Stavros Ntalampiras, Federico Avanzini

arXiv: 1902.05347 · 2019-02-15

## TL;DR

This survey reviews multimodal music information processing techniques, categorizes existing approaches, and discusses future challenges to enhance music computing applications using diverse data modalities.

## Contribution

It provides a comprehensive categorization of multimodal approaches in music information retrieval and highlights key future challenges for the research community.

## Key findings

- Multimodal data enhances music information retrieval accuracy.
- Existing fusion approaches vary in effectiveness and complexity.
- Identified challenges include data integration and modality-specific processing.

## Abstract

Towards improving the performance in various music information processing tasks, recent studies exploit different modalities able to capture diverse aspects of music. Such modalities include audio recordings, symbolic music scores, mid-level representations, motion, and gestural data, video recordings, editorial or cultural tags, lyrics and album cover arts. This paper critically reviews the various approaches adopted in Music Information Processing and Retrieval and highlights how multimodal algorithms can help Music Computing applications. First, we categorize the related literature based on the application they address. Subsequently, we analyze existing information fusion approaches, and we conclude with the set of challenges that Music Information Retrieval and Sound and Music Computing research communities should focus in the next years.

## Full text

_Full body text omitted from this summary view._ Fetch the complete paper as Markdown: https://tomesphere.com/paper/1902.05347/full.md

## Figures

4 figures with captions in the complete paper: https://tomesphere.com/paper/1902.05347/full.md

## References

80 references — full list in the complete paper: https://tomesphere.com/paper/1902.05347/full.md

---
Source: https://tomesphere.com/paper/1902.05347