Model as Loss: A Self-Consistent Training Paradigm

Saisamarth Rajesh Phaye; Milos Cernak; Andrew Harper

arXiv:2505.21156·cs.SD·May 28, 2025

Model as Loss: A Self-Consistent Training Paradigm

Saisamarth Rajesh Phaye, Milos Cernak, Andrew Harper

PDF

Open Access

TL;DR

This paper introduces a novel training paradigm called Model as Loss, which uses the model's own encoder as a loss function to improve speech enhancement by capturing perceptual and task-specific features.

Contribution

The paper proposes a self-consistent training framework that replaces handcrafted or pre-trained feature losses with the model's encoder as a loss, enhancing speech enhancement performance.

Findings

01

Outperforms pre-trained deep feature losses on benchmarks

02

Improves perceptual quality of speech enhancement

03

Generalizes well to in-domain and out-of-domain data

Abstract

Conventional methods for speech enhancement rely on handcrafted loss functions (e.g., time or frequency domain losses) or deep feature losses (e.g., using WavLM or wav2vec), which often fail to capture subtle signal properties essential for optimal performance. To address this, we propose Model as Loss, a novel training paradigm that utilizes the encoder from the same model as a loss function to guide the training. The Model as Loss paradigm leverages the encoder's task-specific feature space, optimizing the decoder to produce output consistent with perceptual and task-relevant characteristics of the clean signal. By using the encoder's learned features as a loss function, this framework enforces self-consistency between the clean reference speech and the enhanced model output. Our approach outperforms pre-trained deep feature losses on standard speech enhancement benchmarks, offering…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech and Audio Processing · Hearing Loss and Rehabilitation · Face recognition and analysis