Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

Vamshi Nallaguntla; Shruti Kshirsagar; Anderson R. Avila

arXiv:2605.03079·cs.SD·May 6, 2026

Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

Vamshi Nallaguntla, Shruti Kshirsagar, Anderson R. Avila

PDF

TL;DR

This paper introduces a phoneme-level analysis framework using self-supervised embeddings to improve detection of emotionally manipulated synthetic speech, enhancing interpretability and robustness.

Contribution

It proposes a novel phoneme-level approach with WavLM embeddings for detecting emotional deepfakes, addressing limitations of previous homogeneous signal methods.

Findings

01

Complex vowels and fricatives show higher divergence in synthetic speech.

02

Phonemes with larger distributional differences are more detectable.

03

Phoneme-level analysis improves interpretability in deepfake detection.

Abstract

Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal and largely overlook its internal phonetic structure, limiting their interpretability in emotionally conditioned settings. In this work, we propose a phoneme-level framework to analyze emotionally manipulated synthetic speech using real and EVC-generated speech under matched emotional conditions with shared transcripts, phoneme-aligned TextGrids, and WavLM-based embeddings. Our results show that phoneme behavior varies across categories, with complex vowels and fricatives exhibiting higher divergence while simpler phonemes remain more stable. Phonemes with larger distributional differences are also found to be more easily detected, consistently across multiple emotions…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.