# SDiFL: Stable Diffusion-Driven Framework for Image Forgery Localization

**Authors:** Yang Su, Shunquan Tan, Jiwu Huang

arXiv: 2508.20182 · 2025-08-29

## TL;DR

SDiFL introduces a novel framework that integrates Stable Diffusion's multi-modal capabilities with image forgery localization, achieving significant accuracy improvements and robustness on diverse datasets.

## Contribution

This work is the first to condition Stable Diffusion's architecture on forgery-related information for improved localization accuracy.

## Key findings

- Achieves up to 12% performance improvement over state-of-the-art methods.
- Effectively localizes forgeries in real-world document and scene images.
- Maintains semantic richness of images while enhancing forgery detection.

## Abstract

Driven by the new generation of multi-modal large models, such as Stable Diffusion (SD), image manipulation technologies have advanced rapidly, posing significant challenges to image forensics. However, existing image forgery localization methods, which heavily rely on labor-intensive and costly annotated data, are struggling to keep pace with these emerging image manipulation technologies. To address these challenges, we are the first to integrate both image generation and powerful perceptual capabilities of SD into an image forensic framework, enabling more efficient and accurate forgery localization. First, we theoretically show that the multi-modal architecture of SD can be conditioned on forgery-related information, enabling the model to inherently output forgery localization results. Then, building on this foundation, we specifically leverage the multimodal framework of Stable DiffusionV3 (SD3) to enhance forgery localization performance.We leverage the multi-modal processing capabilities of SD3 in the latent space by treating image forgery residuals -- high-frequency signals extracted using specific highpass filters -- as an explicit modality. This modality is fused into the latent space during training to enhance forgery localization performance. Notably, our method fully preserves the latent features extracted by SD3, thereby retaining the rich semantic information of the input image. Experimental results show that our framework achieves up to 12% improvements in performance on widely used benchmarking datasets compared to current state-of-the-art image forgery localization models. Encouragingly, the model demonstrates strong performance on forensic tasks involving real-world document forgery images and natural scene forging images, even when such data were entirely unseen during training.

## Full text

_Full body text omitted from this summary view._ Fetch the complete paper as Markdown: https://tomesphere.com/paper/2508.20182/full.md

## Figures

10 figures with captions in the complete paper: https://tomesphere.com/paper/2508.20182/full.md

## References

42 references — full list in the complete paper: https://tomesphere.com/paper/2508.20182/full.md

---
Source: https://tomesphere.com/paper/2508.20182