BitsFusion: 1.99 bits Weight Quantization of Diffusion Model

Yang Sui; Yanyu Li; Anil Kag; Yerlan Idelbayev; Junli Cao; Ju Hu,; Dhritiman Sagar; Bo Yuan; Sergey Tulyakov; Jian Ren

arXiv:2406.04333·cs.CV·October 29, 2024·1 cites

BitsFusion: 1.99 bits Weight Quantization of Diffusion Model

Yang Sui, Yanyu Li, Anil Kag, Yerlan Idelbayev, Junli Cao, Ju Hu,, Dhritiman Sagar, Bo Yuan, Sergey Tulyakov, Jian Ren

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper introduces a novel weight quantization method for diffusion models, reducing model size to 1.99 bits per weight and improving generation quality, enabling efficient deployment on resource-limited devices.

Contribution

The paper presents a new quantization technique that achieves near 2-bit weights for diffusion models, with optimized layer-wise bit assignment and training strategies for better performance.

Findings

01

Model size reduced by 7.9x

02

Generation quality surpasses original models

03

Extensive evaluation confirms effectiveness

Abstract

Diffusion-based image generation models have achieved great success in recent years by showing the capability of synthesizing high-quality content. However, these models contain a huge number of parameters, resulting in a significantly large model size. Saving and transferring them is a major bottleneck for various applications, especially those running on resource-constrained devices. In this work, we develop a novel weight quantization method that quantizes the UNet from Stable Diffusion v1.5 to 1.99 bits, achieving a model with 7.9X smaller size while exhibiting even better generation quality than the original one. Our approach includes several novel techniques, such as assigning optimal bits to each layer, initializing the quantized model for better performance, and improving the training strategy to dramatically reduce quantization error. Furthermore, we extensively evaluate our…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

huggingface/diffusers
jaxOfficial

Videos

BitsFusion: 1.99 bits Weight Quantization of Diffusion Model· slideslive

Taxonomy

TopicsNeural Networks and Applications

MethodsDiffusion