BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Wei Huang; Yangdong Liu; Haotong Qin; Ying Li; Shiming Zhang,; Xianglong Liu; Michele Magno; Xiaojuan Qi

arXiv:2402.04291·cs.LG·June 19, 2024·25 cites

BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang,, Xianglong Liu, Michele Magno, Xiaojuan Qi

PDF

Open Access 1 Repo

TL;DR

BiLLM introduces a novel 1-bit post-training quantization method for large language models, significantly reducing memory and computation needs while maintaining high accuracy, and demonstrating practical efficiency on large models.

Contribution

BiLLM is the first to achieve high-accuracy 1-bit quantization of LLMs using a novel weight selection and binary residual approximation strategy.

Findings

01

Achieves 8.41 perplexity on LLaMA2-70B with 1.08-bit weights.

02

Outperforms state-of-the-art quantization methods for LLMs.

03

Binarizes 7-billion-parameter models within 0.5 hours on a single GPU.

Abstract

Pretrained large language models (LLMs) exhibit exceptional general language processing capabilities but come with significant demands on memory and computational resources. As a powerful compression technology, binarization can extremely reduce model weights to a mere 1 bit, lowering the expensive computation and memory requirements. However, existing quantization techniques fall short of maintaining LLM performance under ultra-low bit-widths. In response to this challenge, we present BiLLM, a groundbreaking 1-bit post-training quantization scheme tailored for pretrained LLMs. Based on the weight distribution of LLMs, BiLLM first identifies and structurally selects salient weights, and minimizes the compression loss through an effective binary residual approximation strategy. Moreover, considering the bell-shaped distribution of the non-salient weights, we propose an optimal splitting…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

aaronhuang-778/billm
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSemantic Web and Ontologies · Natural Language Processing Techniques