Residual Recurrent CRNN for End-to-End Optical Music Recognition on   Monophonic Scores

Aozhi Liu; Lipei Zhang; Yaqi Mei; Baoqiang Han; Zifeng Cai; Zhaohua; Zhu; Jing Xiao

arXiv:2010.13418·cs.CV·August 5, 2021·1 cites

Residual Recurrent CRNN for End-to-End Optical Music Recognition on Monophonic Scores

Aozhi Liu, Lipei Zhang, Yaqi Mei, Baoqiang Han, Zifeng Cai, Zhaohua, Zhu, Jing Xiao

PDF

Open Access

TL;DR

This paper introduces a novel Residual Recurrent CRNN framework that enhances end-to-end optical music recognition accuracy for monophonic scores by better capturing contextual information.

Contribution

It combines residual recurrent convolutional blocks with an encoder-decoder to improve context understanding in music symbol transcription.

Findings

01

Outperforms previous end-to-end CRNN models on CAMERA-PRIMUS dataset.

02

Enhances context information extraction through residual recurrent blocks.

03

Achieves state-of-the-art accuracy in optical music recognition for monophonic scores.

Abstract

One of the challenges of the Optical Music Recognition task is to transcript the symbols of the camera-captured images into digital music notations. Previous end-to-end model which was developed as a Convolutional Recurrent Neural Network does not explore sufficient contextual information from full scales and there is still a large room for improvement. We propose an innovative framework that combines a block of Residual Recurrent Convolutional Neural Network with a recurrent Encoder-Decoder network to map a sequence of monophonic music symbols corresponding to the notations present in the image. The Residual Recurrent Convolutional block can improve the ability of the model to enrich the context information. The experiment results are benchmarked against a publicly available dataset called CAMERA-PRIMUS, which demonstrates that our approach surpass the state-of-the-art end-to-end…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Speech and Audio Processing · Music Technology and Sound Studies