An Analysis on Quantizing Diffusion Transformers

Yuewei Yang; Jialiang Wang; Xiaoliang Dai; Peizhao Zhang; Hongbo Zhang

arXiv:2406.11100·cs.CV·June 18, 2024

An Analysis on Quantizing Diffusion Transformers

Yuewei Yang, Jialiang Wang, Xiaoliang Dai, Peizhao Zhang, Hongbo Zhang

PDF

Open Access

TL;DR

This paper introduces a novel, optimization-free post-training quantization method for diffusion transformers, significantly reducing model size and inference cost without sacrificing performance, demonstrated through image generation tasks.

Contribution

It pioneers an efficient PTQ approach for transformer-only diffusion models, employing single-step activation calibration and group-wise weight quantization without optimization.

Findings

01

Effective low-bit quantization achieved without optimization

02

Reduced inference memory and computational requirements

03

Maintained image generation quality in experiments

Abstract

Diffusion Models (DMs) utilize an iterative denoising process to transform random noise into synthetic data. Initally proposed with a UNet structure, DMs excel at producing images that are virtually indistinguishable with or without conditioned text prompts. Later transformer-only structure is composed with DMs to achieve better performance. Though Latent Diffusion Models (LDMs) reduce the computational requirement by denoising in a latent space, it is extremely expensive to inference images for any operating devices due to the shear volume of parameters and feature sizes. Post Training Quantization (PTQ) offers an immediate remedy for a smaller storage size and more memory-efficient computation during inferencing. Prior works address PTQ of DMs on UNet structures have addressed the challenges in calibrating parameters for both activations and weights via moderate optimization. In this…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNeural Networks and Applications

MethodsDiffusion