VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

Zijing Zhao; Kai Wang; Hao Huang; Ying Hu; Liang He; Jichen Yang

arXiv:2506.16020·cs.SD·June 23, 2025

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

Zijing Zhao, Kai Wang, Hao Huang, Ying Hu, Liang He, Jichen Yang

PDF

Open Access

TL;DR

VS-Singer is a novel vision-guided stereo singing voice synthesis model that integrates spatial cues from images to produce realistic stereo audio with room reverberation, using a unified framework.

Contribution

It introduces a new unified framework combining stereo voice synthesis with visual acoustic matching, utilizing a consistency Schrödinger bridge for efficient one-step generation.

Findings

01

Effectively generates stereo singing voices aligned with scene perspective

02

Utilizes a novel consistency Schrödinger bridge for one-step sample generation

03

Improves audio-visual matching consistency

Abstract

To explore the potential advantages of utilizing spatial cues from images for generating stereo singing voices with room reverberation, we introduce VS-Singer, a vision-guided model designed to produce stereo singing voices with room reverberation from scene images. VS-Singer comprises three modules: firstly, a modal interaction network integrates spatial features into text encoding to create a linguistic representation enriched with spatial information. Secondly, the decoder employs a consistency Schr\"odinger bridge to facilitate one-step sample generation. Moreover, we utilize the SFE module to improve the consistency of audio-visual matching. To our knowledge, this study is the first to combine stereo singing voice synthesis with visual acoustic matching within a unified framework. Experimental results demonstrate that VS-Singer can effectively generate stereo singing voices that…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Face recognition and analysis · Music Technology and Sound Studies