Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts

Michael Kuhlmann; Alexander Werning; Thilo von Neumann; Reinhold Haeb-Umbach

arXiv:2601.21886·eess.AS·January 30, 2026

Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts

Michael Kuhlmann, Alexander Werning, Thilo von Neumann, Reinhold Haeb-Umbach

PDF

Open Access

TL;DR

This paper introduces a method to improve speech quality assessment by regularizing utterance-level predictors with segment-based constraints, enabling better detection of synthesis artifacts and spoofing in speech systems.

Contribution

It proposes a novel regularization technique for utterance-level speech quality models using segment-based constraints, enhancing interpretability and application in artifact detection.

Findings

01

Regularized models show reduced frame-level stochasticity.

02

Frame-level scores correlate with perceived quality in listening tests.

03

Method effectively detects synthesis artifacts in TTS systems.

Abstract

A large number of works view the automatic assessment of speech from an utterance- or system-level perspective. While such approaches are good in judging overall quality, they cannot adequately explain why a certain score was assigned to an utterance. frame-level scores can provide better interpretability, but models predicting them are harder to tune and regularize since no strong targets are available during training. In this work, we show that utterance-level speech quality predictors can be regularized with a segment-based consistency constraint which notably reduces frame-level stochasticity. We then demonstrate two applications involving frame-level scores: The partial spoof scenario and the detection of synthesis artefacts in two state-of-the-art text-to-speech systems. For the latter, we perform listening tests and confirm that listeners rate segments to be of poor quality more…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Speech and Audio Processing · Phonetics and Phonology Research