PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM   Compression

Vladimir Malinovskii; Denis Mazur; Ivan Ilin; Denis Kuznedelev,; Konstantin Burlachenko; Kai Yi; Dan Alistarh; Peter Richtarik

arXiv:2405.14852·cs.LG·May 31, 2024

PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression

Vladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev,, Konstantin Burlachenko, Kai Yi, Dan Alistarh, Peter Richtarik

PDF

Open Access 1 Repo 10 Models 1 Video

TL;DR

This paper introduces PV-Tuning, a novel fine-tuning framework for extreme quantization of large language models, surpassing prior methods in accuracy and efficiency, especially at 1-2 bits per parameter.

Contribution

PV-Tuning is a representation-agnostic, improved fine-tuning approach that outperforms existing methods and guarantees convergence, enabling Pareto-optimal quantization of Llama 2 models at 2 bits.

Findings

01

PV-Tuning outperforms prior quantization techniques on Llama and Mistral models.

02

Achieves Pareto-optimal 2-bit quantization for Llama 2 models.

03

Demonstrates the limitations of straight-through estimators in extreme LLM compression.

Abstract

There has been significant interest in "extreme" compression of large language models (LLMs), i.e., to 1-2 bits per parameter, which allows such models to be executed efficiently on resource-constrained devices. Existing work focused on improved one-shot quantization techniques and weight representations; yet, purely post-training approaches are reaching diminishing returns in terms of the accuracy-vs-bit-width trade-off. State-of-the-art quantization methods such as QuIP# and AQLM include fine-tuning (part of) the compressed parameters over a limited amount of calibration data; however, such fine-tuning techniques over compressed weights often make exclusive use of straight-through estimators (STE), whose performance is not well-understood in this setting. In this work, we question the use of STE for extreme LLM compression, showing that it can be sub-optimal, and perform a systematic…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

vahe1994/aqlm
pytorchOfficial

Models

Videos

PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression· slideslive

Taxonomy

TopicsPhotovoltaic System Optimization Techniques

MethodsLLaMA