Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
Chang Liu, Haomin Zhang, Shiyu Xia, Zihao Chen, Chaofan Ding, Xin Yue, Huizhe Chen, Xinhan Di

TL;DR
This paper introduces the CoP Benchmark Dataset, a comprehensive multimodal benchmark with detailed annotations and evaluation tools, aimed at advancing high-quality video-guided piano music generation.
Contribution
It presents the first open-source, multimodal benchmark with step-by-step Chain-of-Perform guidance for evaluating video-to-piano music generation methods.
Findings
Provides detailed multimodal annotations for precise alignment.
Offers a versatile evaluation framework for various generation tasks.
Includes a publicly available dataset and ongoing leaderboard.
Abstract
Generating high-quality piano audio from video requires precise synchronization between visual cues and musical output, ensuring accurate semantic and temporal alignment.However, existing evaluation datasets do not fully capture the intricate synchronization required for piano music generation. A comprehensive benchmark is essential for two primary reasons: (1) existing metrics fail to reflect the complexity of video-to-piano music interactions, and (2) a dedicated benchmark dataset can provide valuable insights to accelerate progress in high-quality piano music generation. To address these challenges, we introduce the CoP Benchmark Dataset-a fully open-sourced, multimodal benchmark designed specifically for video-guided piano music generation. The proposed Chain-of-Perform (CoP) benchmark offers several compelling features: (1) detailed multimodal annotations, enabling precise semantic…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsMusic Technology and Sound Studies · Computer Graphics and Visualization Techniques · Advanced Vision and Imaging
