An Efficient Framework for Crediting Data Contributors of Diffusion   Models

Chris Lin; Mingyu Lu; Chanwoo Kim; Su-In Lee

arXiv:2407.03153·cs.LG·March 5, 2025

An Efficient Framework for Crediting Data Contributors of Diffusion Models

Chris Lin, Mingyu Lu, Chanwoo Kim, Su-In Lee

PDF

Open Access 1 Video

TL;DR

This paper presents a computationally efficient framework for attributing global properties of diffusion models to data contributors using Shapley values, enabling better incentives and policies for data sharing.

Contribution

It introduces a novel method leveraging model pruning and fine-tuning to efficiently estimate Shapley values for diffusion models, addressing computational challenges.

Findings

01

Accurately identifies important data contributors for diffusion models.

02

Outperforms existing attribution methods in experiments.

03

Effective across different diffusion model use cases.

Abstract

As diffusion models are deployed in real-world settings, and their performance is driven by training data, appraising the contribution of data contributors is crucial to creating incentives for sharing quality data and to implementing policies for data compensation. Depending on the use case, model performance corresponds to various global properties of the distribution learned by a diffusion model (e.g., overall aesthetic quality). Hence, here we address the problem of attributing global properties of diffusion models to data contributors. The Shapley value provides a principled approach to valuation by uniquely satisfying game-theoretic axioms of fairness. However, estimating Shapley values for diffusion models is computationally impractical because it requires retraining on many training data subsets corresponding to different contributors and rerunning inference. We introduce a…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

An Efficient Framework for Crediting Data Contributors of Diffusion Models· slideslive

Taxonomy

TopicsNeural Networks and Applications · Bayesian Modeling and Causal Inference · Constraint Satisfaction and Optimization

MethodsDiffusion · Pruning