Music Style Transfer With Diffusion Model

Hong Huang; Yuyi Wang; Luyao Li; Jun Lin

arXiv:2404.14771·cs.SD·April 24, 2024

Music Style Transfer With Diffusion Model

Hong Huang, Yuyi Wang, Luyao Li, Jun Lin

PDF

Open Access

TL;DR

This paper introduces a diffusion model-based framework for multi-style music transfer that improves audio quality, reduces noise, and enables real-time generation on consumer hardware.

Contribution

It presents a novel diffusion model approach for multi-style music transfer, addressing computational efficiency and audio artifact issues of prior methods.

Findings

01

High-quality multi-style music transfer achieved

02

Real-time audio generation on consumer GPUs demonstrated

03

Reduced noise and artifacts in generated audio

Abstract

Previous studies on music style transfer have mainly focused on one-to-one style conversion, which is relatively limited. When considering the conversion between multiple styles, previous methods required designing multiple modes to disentangle the complex style of the music, resulting in large computational costs and slow audio generation. The existing music style transfer methods generate spectrograms with artifacts, leading to significant noise in the generated audio. To address these issues, this study proposes a music style transfer framework based on diffusion models (DM) and uses spectrogram-based methods to achieve multi-to-multi music style transfer. The GuideDiff method is used to restore spectrograms to high-fidelity audio, accelerating audio generation speed and reducing noise in the generated audio. Experimental results show that our model has good performance in multi-mode…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMusic and Audio Processing · Speech and Audio Processing

MethodsSPEED: Separable Pyramidal Pooling EncodEr-Decoder for Real-Time Monocular Depth Estimation on Low-Resource Settings · Diffusion