PSVRF: Learning to restore Pitch-Shifted Voice without reference

Yangfu Li; Xiaodan Lin; and Jiaxin Yang

arXiv:2210.02731·cs.SD·March 14, 2023

PSVRF: Learning to restore Pitch-Shifted Voice without reference

Yangfu Li, Xiaodan Lin, and Jiaxin Yang

PDF

Open Access 1 Repo

TL;DR

This paper introduces PSVRF, a no-reference method for restoring pitch-shifted voices, significantly improving ASV system robustness against pitch-scaling attacks without needing the original voice as a reference.

Contribution

PSVRF is the first no-reference approach that effectively restores pitch-shifted voices, outperforming existing reference-based methods in quality and robustness.

Findings

01

PSVRF successfully restores pitch-shifted voices across various techniques.

02

It enhances the robustness of ASV systems against pitch-scaling attacks.

03

PSVRF outperforms state-of-the-art reference-based approaches.

Abstract

Pitch scaling algorithms have a significant impact on the security of Automatic Speaker Verification (ASV) systems. Although numerous anti-spoofing algorithms have been proposed to identify the pitch-shifted voice and even restore it to the original version, they either have poor performance or require the original voice as a reference, limiting the prospects of applications. In this paper, we propose a no-reference approach termed PSVRF $^{1}$ for high-quality restoration of pitch-shifted voice. Experiments on AISHELL-1 and AISHELL-3 demonstrate that PSVRF can restore the voice disguised by various pitch-scaling techniques, which obviously enhances the robustness of ASV systems to pitch-scaling attacks. Furthermore, the performance of PSVRF even surpasses that of the state-of-the-art reference-based approach.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

ychenl/pssrf
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Speech and Audio Processing · Music and Audio Processing