On Statistical Rates and Provably Efficient Criteria of Latent Diffusion   Transformers (DiTs)

Jerry Yao-Chieh Hu; Weimin Wu; Zhao Song; Han Liu

arXiv:2407.01079·stat.ML·November 1, 2024

On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)

Jerry Yao-Chieh Hu, Weimin Wu, Zhao Song, Han Liu

PDF

Open Access 1 Video

TL;DR

This paper analyzes the statistical approximation capabilities and computational efficiency of latent Diffusion Transformers (DiTs) under low-dimensional assumptions, providing bounds, criteria, and algorithms for faster inference and training.

Contribution

It offers new theoretical bounds on DiTs score function approximation, distribution recovery, and proposes efficient algorithms for inference and training leveraging low-rank structures.

Findings

01

Sub-linear approximation error bound in latent space dimension

02

Almost-linear time inference algorithms for latent DiTs

03

Low-rank gradient approximations enable faster training

Abstract

We investigate the statistical and computational limits of latent Diffusion Transformers (DiTs) under the low-dimensional linear latent space assumption. Statistically, we study the universal approximation and sample complexity of the DiTs score function, as well as the distribution recovery property of the initial data. Specifically, under mild data assumptions, we derive an approximation error bound for the score network of latent DiTs, which is sub-linear in the latent space dimension. Additionally, we derive the corresponding sample complexity bound and show that the data distribution generated from the estimated score function converges toward a proximate area of the original one. Computationally, we characterize the hardness of both forward inference and backward computation of latent DiTs, assuming the Strong Exponential Time Hypothesis (SETH). For forward inference, we identify…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)· slideslive

Taxonomy

TopicsNeural Networks and Applications · Speech Recognition and Synthesis

MethodsDiffusion